Repository navigation
About nextQuery #99
Description
Activity
When I run a query, I'm only interested in getting the entities that match my query back.
I can add a special case for limit queries. If there is a limit set and I haven't retrieved all the pages, keep doing it until the end and call callback. Does it sound good?
Sure, but what happens if I don't set a limit and do a query? Do I have to deal with the recursive query running still?
I'm just trying to understand why we want to give the developer control beyond the customization of the query. For example, if they don't set a limit, and their query yields a huge amount of results, what are the reasons they wouldn't want to keep running
nextQueryto get the full set of data they queried for?I'm just trying to understand why we want to give the developer control beyond the customization of the query
Typical batch jobs.
- Retrieve some records
- Process them
- Retrieve more or schedule to retrieve more (User usually requires some level of control to fetch more data)
@pcostell, is it OK to continuously keep retrieving new pages in the background if limit is not met yet? We do pagination already to meet some performance constraints of the the backend, would it be harsh on the API if we implement auto pagination for the limit queries at the client level?
You definitely can, although you might cause unintended hurt if people start setting the limit to an arbitrarily high number but don't necessarily want that limit (it happens).
One of the key things to remember, especially if you look at the App Engine client libraries, is that Cloud Datastore is much better about giving you all the data that you want. In App Engine, the datastore has two explicit calls for queries,
RunQueryandNext.Nextallows the query to be continued, and is used pretty extensively for prefetching throughout the App Engine client libraries. However, the size of the result set returned fromRunQueryis much smaller than that returned by Cloud Datastore. In particular, if you are writing a latency-sensitive application (i.e. user-facing stuff) and Cloud Datastore doesn't return the number of results that you wanted, the user should probably only display those results since getting a result means you're hitting a large size or time limit.In the case of a user wanting to do a query and get all the results then process them all together, it may be best to let them hit the point where Cloud Datastore returns and force them to continue manually if that's really what they want.
Typical batch jobs.
- Retrieve some records
- Process them
- Retrieve more or schedule to retrieve more (User usually requires some level of control to fetch more data)
I think here is really where this would be great. It definitely seems like a useful feature for users trying to do large workloads. In particular, doing something where the next batch is being retrieved as the user processes the current batch.
Note that this also might be a good place for the user to use a data processing framework because they can run queries in parallel (although the supported feature set of those queries is limited). We use a special scattered property to split their entire data into chunks which can be processed in parallel.
To answer your specific question, the Datastore API shouldn't have problems if you implement auto-pagination, however specifically for slow queries it may be a poor experience for users if they don't regain control until all results have been fetched.
Thanks for the detailed reply, @pcostell! It sounds like we should keep the API the same with regards to always having to manually invoke
nextQuery. It would likely result in confusion and/or misuse if we had a work-around that changed the behavior of the callback. Feel free to re-open if anyone feels differently 👍I think we should keep providing
runQueryas it is, but can provide helper iterators on top of that. Let's make a release, as people start to use the client, we will see patterns of replication. We might prefer to address them on a different layer and keep havingrunQueryas the most granular option for those who need performance.- addedapi: datastoreIssues related to the Datastore API.Issues related to the Datastore API.
on Feb 2, 2015 - addedtriage meI really want to be triaged.I really want to be triaged.🚨This issue needs some love.This issue needs some love.
on Apr 6, 2020 - added a commit that references this issue
on Aug 22, 2022 - added a commit that references this issue
on Sep 12, 2022 69 remaining items
- added a commit that references this issue
on Feb 26, 2026 - added a commit that references this issue
on Mar 5, 2026 - added a commit that references this issue
on Mar 11, 2026 - added a commit that references this issue
on Mar 12, 2026 - added a commit that references this issue
on Mar 18, 2026 - added 2 commits that reference this issue
on Mar 23, 2026 - added a commit that references this issue
on Mar 27, 2026 - added a commit that references this issue
on May 5, 2026
Can we talk about this again? 😊
To use an example from datastore (but equally applies to the other service calls as well):
Can you show me an example of how this is meant to be used? I'll take a guess, which I'll try to use to demonstrate the trouble I'm having with it:
There won't be anything I want to do in my first callback that I won't want to do in my second callback, or I would have run a different query, right? How would I manage doing something different on the second callback than the first, and something even different on the third, etc? That would be very unmanageable, and the reason someone would write 3 different queries if they needed 3 types of results.
When I run a query, I'm only interested in getting the entities that match my query back. Even if my query is loose enough to send me back 100,000 results, it would be my fault that I didn't restrict the limit in my query or specify strict enough matching rules in the query.
I know I asked about this before, but an "explain it like I'm 5" explanation would be welcome for why we can't just only call the callback once all results are in. :)