Skip to content

About nextQuery #99

Description

@stephenplusplus

Can we talk about this again? 😊

To use an example from datastore (but equally applies to the other service calls as well):

var query = ds.createQuery('Company').limit(5);
ds.runQuery(q, function(err, entities, nextQuery) {
    // nextQuery is not null if there are more results.
    if (nextQuery) {
        ds.runQuery(nextQuery, callback);
    }
});

Can you show me an example of how this is meant to be used? I'll take a guess, which I'll try to use to demonstrate the trouble I'm having with it:

var results = [];
function querySuccess(err, entities, nextQuery) {
  results.push(entities);

  if (nextQuery) {
    ds.runQuery(nextQuery, querySuccess);
  } else {
    nextThingIWantToDo();
  }
}

var q = ds.createQuery('Company').limit(5);
ds.runQuery(q, querySuccess);

There won't be anything I want to do in my first callback that I won't want to do in my second callback, or I would have run a different query, right? How would I manage doing something different on the second callback than the first, and something even different on the third, etc? That would be very unmanageable, and the reason someone would write 3 different queries if they needed 3 types of results.

When I run a query, I'm only interested in getting the entities that match my query back. Even if my query is loose enough to send me back 100,000 results, it would be my fault that I didn't restrict the limit in my query or specify strict enough matching rules in the query.

I know I asked about this before, but an "explain it like I'm 5" explanation would be welcome for why we can't just only call the callback once all results are in. :)

Activity

  1. rakyll commented on Aug 7, 2014

    @rakyll
    Contributor

    When I run a query, I'm only interested in getting the entities that match my query back.

    I can add a special case for limit queries. If there is a limit set and I haven't retrieved all the pages, keep doing it until the end and call callback. Does it sound good?

  2. stephenplusplus commented on Aug 7, 2014

    @stephenplusplus
    ContributorAuthor

    Sure, but what happens if I don't set a limit and do a query? Do I have to deal with the recursive query running still?

    I'm just trying to understand why we want to give the developer control beyond the customization of the query. For example, if they don't set a limit, and their query yields a huge amount of results, what are the reasons they wouldn't want to keep running nextQuery to get the full set of data they queried for?

  3. rakyll commented on Aug 7, 2014

    @rakyll
    Contributor

    I'm just trying to understand why we want to give the developer control beyond the customization of the query

    Typical batch jobs.

    • Retrieve some records
    • Process them
    • Retrieve more or schedule to retrieve more (User usually requires some level of control to fetch more data)

    @pcostell, is it OK to continuously keep retrieving new pages in the background if limit is not met yet? We do pagination already to meet some performance constraints of the the backend, would it be harsh on the API if we implement auto pagination for the limit queries at the client level?

  4. pcostell commented on Aug 7, 2014

    @pcostell
    Contributor

    You definitely can, although you might cause unintended hurt if people start setting the limit to an arbitrarily high number but don't necessarily want that limit (it happens).

    One of the key things to remember, especially if you look at the App Engine client libraries, is that Cloud Datastore is much better about giving you all the data that you want. In App Engine, the datastore has two explicit calls for queries, RunQuery and Next. Next allows the query to be continued, and is used pretty extensively for prefetching throughout the App Engine client libraries. However, the size of the result set returned from RunQuery is much smaller than that returned by Cloud Datastore. In particular, if you are writing a latency-sensitive application (i.e. user-facing stuff) and Cloud Datastore doesn't return the number of results that you wanted, the user should probably only display those results since getting a result means you're hitting a large size or time limit.

    In the case of a user wanting to do a query and get all the results then process them all together, it may be best to let them hit the point where Cloud Datastore returns and force them to continue manually if that's really what they want.

    Typical batch jobs.

    • Retrieve some records
    • Process them
    • Retrieve more or schedule to retrieve more (User usually requires some level of control to fetch more data)

    I think here is really where this would be great. It definitely seems like a useful feature for users trying to do large workloads. In particular, doing something where the next batch is being retrieved as the user processes the current batch.

    Note that this also might be a good place for the user to use a data processing framework because they can run queries in parallel (although the supported feature set of those queries is limited). We use a special scattered property to split their entire data into chunks which can be processed in parallel.

    To answer your specific question, the Datastore API shouldn't have problems if you implement auto-pagination, however specifically for slow queries it may be a poor experience for users if they don't regain control until all results have been fetched.

  5. stephenplusplus commented on Aug 7, 2014

    @stephenplusplus
    ContributorAuthor

    Thanks for the detailed reply, @pcostell! It sounds like we should keep the API the same with regards to always having to manually invoke nextQuery. It would likely result in confusion and/or misuse if we had a work-around that changed the behavior of the callback. Feel free to re-open if anyone feels differently 👍

  6. rakyll commented on Aug 7, 2014

    @rakyll
    Contributor

    I think we should keep providing runQuery as it is, but can provide helper iterators on top of that. Let's make a release, as people start to use the client, we will see patterns of replication. We might prefer to address them on a different layer and keep having runQuery as the most granular option for those who need performance.

  7. added this to the Datastore Stable milestone on Feb 2, 2015
  8. 69 remaining items

  9. added a commit that references this issue on Feb 26, 2026
  10. added a commit that references this issue on Mar 5, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

🚨This issue needs some love.api: datastoreIssues related to the Datastore API.triage meI really want to be triaged.

Type

No type

Projects

No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions