Repository navigation
Distributed Version #23
Description
Activity
I think that's the point - Google hasn't open sourced "scalable" version :)
Thanks for the question! To reiterate what I said here, we are working on making a distributed implementation available, it's currently not in the initial release. Please stay tuned, and take a look at the cifar multi-gpu tutorial for a flavor of how we handle multiple 'devices': http://tensorflow.org/tutorials/deep_cnn/index.md
Reacted by Thomas Wood, Andrei Simionescu and Rodrigo Menezeswould appreciate any insight on the availability of the distributed version. Is the distributed code that is being worked on in github? That is one place where some of us who are interested can contribute
👍
Hello,
After reading these plans and ideas, I'm somewhat surprised. According to http://static.googleusercontent.com/media/research.google.com/en//people/jeff/BayLearn2015.pdf, both data and model parallel are needed to train large and powerful models quickly. BTW, GPUs transferring data takes time as described in http://tensorflow.org/tutorials/deep_cnn/index.md. Then, how it's possible to efficiently support both model parallelism and heterogeneous multi-devices (of a single node) on distributed cluster? Could you please roughly explain how different it from DistBelief?
Thanks!
P.S., GPU acceleration also could be limited by model partition strategies.
Our current internal distributed extensions are somewhat entangled with Google internal infrastructure, which is why we released the single-machine version first. The code is not yet in GitHub, because it has dependencies on other parts of the Google code base at the moment, most of which have been trimmed, but there are some remaining ones.
We realize that distributed support is really important, and it's one of the top features we're prioritizing at the moment.
Reacted by Guancheng (G.C.) Chen, Eric Ma, GitHubLeakedPAN, GitHubLeakedMyautsai, Choco, BogusLogin, Kai(luo) Wang, Moa Raji, easyscheduler, Nawaid Shamim, intoraw and 8 moreReacted by Thomas Wood, GitHubLeakedPAN, GitHubLeakedMyautsai, Mayoor Rao, tobe, Nawaid Shamim, Chenglei Niu, Victor Ziliang Peng, RAIN, Yi Wang and Hao TangAwesome.
After reading the whitepaper, I just realized that large neural network model can be partitioned into sub-graphs by layer (horizontal partitioning) and executed in a serial way.
One thing not clear is the performance for fully connected network on multi node equipped with GPUs cluster ..
In theory, something like Dask could be layered on top for handling this - at least for the Python front-end.
Any update on timeline?
Dask looks interesting project but the drawback of the blocking algorithm is that it's not memory optimal. Since a large amount of memory is required for fully-connected layers, I was thought that Pregel-like model parallelism on CPUs w/ vertical partitioning is more attractive for fully connected layers (blocking mat-mult on GPU also appears to me slow and memory demanding). Of course, I maybe wrong but that's why I launched Apache Horn project recently. Since layers can be pipelined, I hope we can collaborate each projects in complementary way.
I don't know if you can give us some anticipation on the framework choice but..
@bhack I wonder what their Java/Scala interface looks like...
61 remaining items
- added a commit that references this issue
on Dec 6, 2019 - added a commit that references this issue
on Feb 1, 2021 - added 5 commits that reference this issue
on Apr 9, 2025 - added 5 commits that reference this issue
on Jul 28, 2025 - added a commit that references this issue
on Nov 26, 2025
Is there any distributed version of TensorFlow that could work on multiple machines?
-Minjie