Repository navigation
mapSync on Workers - and possibly on the main thread #2217
Description
Activity
Some customers are very interested in this, so I'd like to put it on the meeting agenda.
Without synchronous access to data, we are running into problems with TFJS on WebGPU backend (See tensorflow/tfjs#1595 which suggests that dataSync might never be supported for WebGPU).
Our rendering strategy calls TFJS as a part of the render loop (since we might want to run an augmented video feed through a neural network) and requires access to network results synchronously for further processing. Yielding the render loop to wait for TFJS results would require some significant refactoring and also introduces an overhead to preserve the renderer state while waiting for the results. Otherwise, we need to build some hacks like trying to reuse latest available results and hope that they are applicable to the frame being processed.
When rendering is happening on a worker thread, there shouldn't be an issue in synchronously waiting for GPU to finish enqueued tasks (and thus support
dataSyncin TFJS) since the main thread can continue processing its event loop without stalls.Reacted by Lv YitianWanted to raise our hand as someone who needs blocking operations on workers as well! ✋ (we're compute-focused and don't have concepts of frames or clearly delineated times when we could yield to the browser, and run all our code in workers)
This class of issues is currently our biggest blocker with using WebGPU in the browser - we've only been able to get things working by using
wgpuDevicePoll(..., /*force_wait=*/true)available in wgpu-native to simulate this behavior - but even that's not great.We do also need onSubmittedWorkDone so long as it's not possible to receive a callback for a submission without first yielding to the browser - we're running our entire program in a worker thread without yielding at any point (as would practically anyone coming from webassembly). In our ideal situation this does not imply blocking but that the callback could be issued from another thread such that we can signal a semaphore that a blocking thread may be waiting on. We really would prefer there to be native semaphores/fences in WebGPU but barring that we need to emulate them and the minimum to do so would be a worker thread we kept around to wait for completion events we used to signal potential waiter threads. Something like:
std::thread event_processing_thread([&]() { while (true) { // we don't want this to be a spin-loop - we're really just telling the API we'd // like callbacks to be made from this thread when there are any pending wgpuInstanceProcessEvents(/*blocking=*/true); } }); FenceFutex fence; wgpuQueueOnSubmittedWorkDone(..., [](WGPUQueueWorkDoneStatus status, void *userdata) { fence.signal(); }); wgpuQueueSubmit(....); // can do other work here, non-blocking // maybe wait - if already fired then this won't block and otherwise it'll wait for the event // to be processed on the thread and signal the fence fence.wait();
An interesting discussion in the Dawn matrix room with @austinEng where he had a really nice idea and I wanted to document here:
There may be a single primitive that solves this for javascript/native and workers: allow async operations to take a pointer (element in a SharedArrayBuffer) and a value to have the WebGPU implementation write to that pointer when the callback completes. Browser environments could then use
Atomics.wait(directly in javascript or via futex emulation in webassembly from emscripten) and native environments could use their system futex/WaitOnAddress/etc. Since browsers all have to implementAtomics.waitanyway (if they support workers) there shouldn't be much additional infra required here, and users can easily interact with this mechanism themselves viaAtomics.notify.Something like:
WGPU_EXPORT void wgpuBufferMapAsync2(WGPUBuffer buffer, WGPUMapModeFlags mode, size_t offset, size_t size, void* addr, uint32_t value, WGPUBufferMapAsyncStatus* status);
A synchronous version can then be written easily by users:
bool wgpuBufferMapSync(WGPUBuffer buffer, WGPUMapModeFlags mode, size_t offset, size_t size) { WGPUBufferMapAsyncStatus status; uint32_t futex = 0; wgpuBufferMapAsync2(buffer, mode, offset, size, &futex, 1, &status); // could do other work here, switch to another thread/worker, etc - here just blocking to emulate sync behavior futexWait(&futex, 1); return checkStatusSuccess(status); // could be device loss, etc }
If all the async methods (
wgpuDeviceCreateComputePipelineAsyncandwgpuQueueSubmit) did this then synchronous versions could be written, or more importantly semi-synchronous versions: an application or framework is free to use multiple threads/workers to coordinate work, watch for signals to perform fencing (always be running one frame ahead, etc) without needing spin-loops. In the browser the main thread could submit work and workers could wait on it without dealing with the transfer of promises or other higher-level constructs while still allowing the main thread to poll for completion.The other advantage of this approach is that it solves the multiple codebase issue: if two frameworks are used in the same application - even if one is written in javascript and the other compiled via webassembly - there's no need for complex scheduling coordination or cross-language interop goo or global WGPUInstance/WGPUDevice waits as everything is scoped.
If we had this I think all our concerns around synchronization would be solvable in user code (we'd be able to perform async compilation, async and overlapped submission, and sync mapping where required) and there'd be little additional needed in the WebGPU API.
This could be emulated with a spin-loop as in my previous example (callbacks take userdata of the addr/value and perform the write themselves), but not needing to have user threads/workers spun up to do this - especially if there was one such thread/worker per framework in an application - would be a real win for both ease of use and resource consumption.
@kainino0x helped clarify that this could be seen as the way to do async when you don't have a top-level event loop - the callbacks would work for main thread javascript or a worker that was yielding, and this would cover everything else
Reacted by Kai NinomiyaAn unfortunate thing about using atomics is that they require SharedArrayBuffer which requires COOP+COEP: https://developer.mozilla.org/en-US/docs/Web/JavaScript/Reference/Global_Objects/SharedArrayBuffer#security_requirements
If SAB were always available (and COOP+COEP just enabled sharing it between threads),
or if atomics were allowed on non-shared arraybuffers,then this would work better.Darn, good point :/
This signal style would be most useful in native or worker contexts with SharedArrayBuffer available (as you're not compiling pthreads code to wasm unless you have them). I think in the case of workers with private memory a blocking tick method that issued the callbacks on the worker that issues the tick would cover the same use cases, though I think the signal style even if not waitable could be useful for ergonomics. I'm imagining the case of dispatching a lot of work and then mapping several of the results: that becomes nice straight-line code of submit->map->map->map->loop on tick until all/any ready->use mapped resources. With the pure callback approach that'd need the user to do more juggling (and it scales with the number of mapped resources, and with each part of the code mapping the resources, etc). Not as efficient as a futex for low-latency wakes but at the point where you have a single thread in isolation you probably don't care much about that. I think even the SharedArrayBuffer method would need some kind of flush method unless the implementations flush automatically, so this may all require the same primitives regardless of approach - it's just if you don't have Atomics.wait you can't block on the signals and deadlock yourself.
A possible direction on this issue is that we just give up and start allowing blocking readbacks in general, even on the main thread. I know we're operating under strong architectural guidance that we don't do this. But in practice it's just forcing people to do terrible workarounds with canvases (reshaping/re-encoding data into 2D rgba8unorm data, writing it into a canvas, and then using existing synchronous readback APIs like toDataURL to read them). Maybe we should get out of the business of trying to force people do to the "right thing" when the wrong thing is already possible, and just provide the primitives. The performance consequences of synchronous readback are not that bad (compared with e.g. synchronous XHR).
Reacted by Ben Vanik, Lv Yitian, charlie roberts, Vladimir Kruzhkov and TÖRÖK AttilaAs someone who participated in the WebGL 1 discussions about the same kind of functionality when trying to do compute/media work there I'd agree with most of that - the abominations we had to create still make me sad, but no lack of functionality eliminated our need to create them. Your example is actually one I had to do as "render this canvas to an animated gif" was a product requirement and I couldn't just say "well reading pixels back is really hard in WebGL" - I just had to build a convincing progress spinner because it took so long :P
What we learned then is that there's definite implementation complexity that comes from this kind of stuff but it's really a tradeoff of what's difficult to do in an implementation vs. what's impossible to do on top of the implementation. Needing users to do some extra work to get the precise behavior they want - knowing each user will want a slightly different behavior - works best when the user is actually able to do that work. IME as long as there is a primitive that allows for asynchronously observable multi-threaded semaphores everything else can be built on top of it - blocking sync, callback-style async, pure polling, or mixed futex-based sync/async. If there's only async callbacks with event-loop driven flushes or synchronous blocking operations then the others can't be (practically) implemented by users and there will always be tension (or whining from people like me :). Requiring the performance-destroying use of asyncify when targeting WebGPU or spinning up multiple workers whose only job is to sit and block on waits both fall into that abomination category (this also relates to discussions around efficient timeline-ordered readback via a hypothetical GPUQueue.readBuffer) - people will still need to do those kind of things regardless of whether the implementations make it easy or efficient, and it'd be in the end-user's (web users, etc) best interest for them to be able to be efficient (fewer workers, less large wired allocations, etc).
67 remaining items
@Morglod WebGPU has
GPUQueue.writeBufferandGPUQueue.writeTexturealready. There isn't an equivalent ofgl.getBufferSubDatathough.Reacted by mwyrzykowskiReacted by Vladimir KruzhkovIndeed we already have synchronous writes via above mentioned
writeTextureandwriteBuffer.Can we better understand the use case for synchronous buffer reading? That is always going to be slow as it blocks the GPU pipeline and removing an async dispatch should be measured but is likely a micro perf gain as @kainino0x notes in #2217 (comment)
Indeed we already have synchronous writes via above mentioned
writeTextureandwriteBuffer.Can we better understand the use case for synchronous buffer reading? That is always going to be slow as it blocks the GPU pipeline and removing an async dispatch should be measured but is likely a micro perf gain as @ kainino0x notes in #2217 (comment)
Beyond the speed issue, a ton of people have dropped into the thread detailing their use cases (including a Figma engineer who had to resort to an ugly hack to ship WebGPU!)
Inserting event-loop yield points into your code isn't just something you can do without thinking. Everything that needs to become async has to think about reentrancy:
Of course the rest of the code often needs to be adapted to handle the asynchrony because it might try to do something in the meantime, and there's a development cost to that.
"A development cost" is underselling things a fair bit. As pointed out in the
Atomics.waitGitHub thread, there is no programming model for developing applications where every function is re-entrant. The Rust standard library itself, as well as Emscripten, are forced to busy-wait because the Web platform folks have deemed it not "best practices" to yield from the main thread. If Emscripten can't figure it out, who can?It's very easy to introduce hard-to-trigger and hard-to-test race conditions with async. And the guts of your fancy rendering or GPGPU engine is the worst possible place to force an event loop yield.
@juj has some good commentary on this point.
Reacted by StanI will say that it's quite frustrating to have several people, over the course of years, drop into this thread detailing their use cases, and still have spec maintainers and implementers asking "can you explain your use case?" I'm sure there's a lot of discussion going on internally, but to outside observers who have to figure out how to actually use these APIs, it's quite opaque.
@mwyrzykowski Well, same as for reading through mapping, but without overhead of mapping. Its per-frame gpu computations.
When you want to have SYNC READ RIGHT NOW, with current api, you should do a lot of unnecessary stuff with unnecessary overhead.
To be honest it feels like now you are working on inventing new excuses rather than on solving this issue.
Reacted by valadaptive@mwyrzykowski Well, same as for reading through mapping, but without overhead of mapping. Its per-frame gpu computations.
@kainino0x mentioned it in #2217 (comment) but in this context the words 'read' and 'mapping' mean the same.
To help illustrate the difference, the current API is implemented something like (very high level pseudo-code):
mapAsync(buffer) { gpuDriver.callbackWhenAllWorkOnResourceFinishes(buffer, () => { fulfill JS promise with contents of 'buffer' }); }it sounds like there is a request for:
mapSync(buffer) { semaphore s; gpuDriver.callbackWhenAllWorkOnResourceFinishes(buffer, () => { synchronously return from JS function with contents of 'buffer', which is possible due to waiting on a semaphore s.signal(); }); s.wait(); }as previously pointed out this will only potentially mitigate a thread switch.
If I'm understanding correctly, it sounds like making sync JS / wasm to become asynchronous is more than a non-trivial effort?
I think we all legitimately want to understand how to help solve this, given the demand :)
Is there need for this outside of workers? For workers, this seems completely reasonable and quite simple for UAs to implement as well.
If this is really desired on the main thread, we would need to do some more thinking due to blocking of the main thread.
I'd love to not need to lock this thread. Please stay civil and assume good faith effort from all parties involved. The request for more info is to 1) figure out if we can find another alternative (sync waiting is a no-no in guidelines for new Web APIs) 2) which workers (or main thread) need this.
Also don't underestimate the effort that this requires to spec and implement. At least in Chromium making JS able to synchronously wait on stuff for WebGPU in GPU process side is a complex change, my guess is 1-2 months of engineering. That's even before any implementation of a
mapSync. There's also a matter of prioritization especially at times where Firefox and WebKit a reaching a first shippable version of WebGPU.Reacted by mwyrzykowskiReacted by Vladimir Kruzhkovsync waiting is a no-no in guidelines for new Web APIs
I think this, fundamentally, is the issue that my project has with this. Existing web APIs, notably WebGL, allow for this, and so existing applications require this. If we want to start switching to WebGPU for more capabilities, performance or whatever - we need to at least have the same capabilities that WebGL allowed for, which means a sync readback of buffers.
For context, our application is an emulator for Flash (.swf) files, which can be wholly emulated right now using webgl. Flash allowed, in fact encouraged, sync readback and therefor we need to allow for it too - which means switching to webgpu is off the table for us, despite in theory being an upgrade in every other way. Our use case then isn't strictly about mapping a buffer sync, but doing a single readback sync.
It's a very good thing to not block the main thread or stall the gpu in new applications, I completely agree - but sometimes it is needed, and we should not dismiss decades of history that used it because by todays standards they were doing something suboptimal.
Reacted by TÖRÖK Attila, Vladimir Kruzhkov, valadaptive, charlie roberts, Alex Ringlein, Kai Ninomiya, moirj15, Ambrose Robinson, nosamu, Daniel Jacobs and 3 moreI've spent some time trying out the recent experimental mapSync in workers support in Chrome and wanted to give some feedback on it.
MapSync (even only in a worker) would have been very useful to us if it had existed. In our case, we have some C++ code that we want to adapt to compile with emscripten and render with WebGPU. Since this code originally ran on desktop, it has the assumption that synchronous readback from the GPU is possible baked into it. This assumption can be a lot of work to refactor out. In our case it was mostly around GPU-based picking. With no way of doing synchronous mapping, our only option was to invest significant effort in a refactor, or fallback to CPU picking. CPU picking is also synchronous, blocking, and can be slow. Having the possibility of using mapSync in a worker context would have saved us from this and made it easier to move forward.
One note on the API: In our case, there are multiple buffers that need to be mapped back before we can proceed. If using
mapAsync, it's possible to make all themapAsyncrequests at once and then do something with all the results when available usingPromise.all(). UsingmapSync, you can't make the next mapping request until the first one resolves.Roughly,
// Faster, parallel requests await Promise.all([buffer1.mapAsync(), buffer2.mapAsync(), buffer3.mapAsync(), buffer4.mapAsync()]) // Slower, sequential requests await buffer1.mapAsync() await buffer2.mapAsync() await buffer3.mapAsync() await buffer4.mapAsync() // About the same as sequential. No option for issuing all requests upfront and then a blocking wait buffer1.mapSync() buffer2.mapSync() buffer3.mapSync() buffer4.mapSync()
I'm unsure how it all works internally, but if it's faster to know upfront about the mapping requests before waiting on the GPU, then being able batch issue the mapping requests would be helpful for our use case. There may be other solutions, like writing all the data into 1 big buffer instead of 4 smaller ones, but I wanted to put this out there.
// A questionable (but helpful) API. Equivalent to the await Promise.all([...]) above, but sync const allResults = GPUBuffer.batchMapSync([buffer1, buffer2, buffer3, buffer4])
Even as the API is, this still would have been very, very helpful to have had.
Reacted by Mauricio Vives and Vladimir KruzhkovDid you consider refactoring to use slightly old picking data?
original picking code
- render to
pickingData - synchronously map
pickingData - synchronously use
pickingData
new picking code
- have
pickingDataandpendingPickingData - render to
pendingPickingData - asynchronously map
pendingPickingData.then() { [pendingPickingData, pickingData] = [pickingData, pendingPickingData]; // swap refs }); - synchronously use
pickingData
This leaves your main loop synchronous and all that happens is the data is a moment off. Bonus, you don't stall the rendering pipeline with a sync map so your app feels more responsive. It's also often a very minor refactor.
- render to
Yes, that is basically the strategy we ended up with for most picking since we often want to cache the pick data for the frame for faster lookups if the view hasn't changed. We would do this even if it was possible to sync map but there is a complication if you can only do async mapping. There are also some picking paths like deep picking which hits deeper than the surface where we can't cache everything.
For caching the pickingData, there is a window of time where you know that pickingData is out of date and that a request is in flight to get the latest pickingData. If a user action that requires picking data (like a click) happens in that window then what do you do? You know that the data you have is wrong so if you base some action on that then it might be wrong. You have a few options:
- Process the event with out of date pickData. You might hit the wrong thing which isn't good depending on what you're trying to do.
- Just return no-hits since you can't know if the hit will be accurate. Might be fine for a mousemove where another event is likely to come in but not great for something like a click
- You can delay processing the action until the latest data is ready. That means you have to queue up input events (or anything that might need a hit test) and delay processing them if you ever have out of date pick data, which is a larger change.
- You can fallback to CPU hit testing. This also blocks and is slower than mapSync. It's also more complicated since you now have 2 different picking mechanisms to keep identical.
For the third option, this is basically the same situation you end up in with sync mapping except with sync mapping a user action will only wait for the data if it's needed (when the mapping is triggered). If you can only do async mapping then you have to know before processing the user action whether or not it will need the pickingData, so that you can decide whether to delay it or not. Otherwise you have to delay all actions until pickData is ready
In practice this window of time may be quite small, but you can get into situations where one action will invalidate the pick data with a camera move, the next action will definitely need picking data, and these will happen very fast. If these bugs do happen, they're a nightmare to track down. Having the synchronous path available can give much more confidence that the ported code is working properly and then we can work towards the async path from there
In principle, AFAIU, synchronous blocking APIs are only problematic on the main thread. Web workers can block on operations without really causing any problems - especially since our operations should never block for very long (unlike network operations for example, although synchronous XHR is not a good example because it's old).
We could mirror some APIs into synchronous versions with
[Exposed=DedicatedWorker].Definitely useful - I'd start with just this one:
EDIT: There are other possible entry points but let's ignore them for now
Maybe useful but possibly not worth the complexity:
Most likely not needed:
Synchronous map (mapSync) is known to be particularly handy. For example, TensorFlow.js can implement its
dataSync()method to synchronously read data back from a gpu-backed tensor (even if only available on workers).Finally, there's no avoiding the fact that poorer forms of synchronous readback are always going to be available (e.g. synchronous canvas readback) so we should at least consider supporting mapSync on workers where it's OK in principle.
EDIT:
Originally posted by @kainino0x in #2217 (comment)