Skip to content

Distributed Version #23

Description

@jermainewang

Is there any distributed version of TensorFlow that could work on multiple machines?

-Minjie

Activity

  1. rusenask commented on Nov 9, 2015

    @rusenask

    I think that's the point - Google hasn't open sourced "scalable" version :)

  2. vrv commented on Nov 9, 2015

    @vrv

    Thanks for the question! To reiterate what I said here, we are working on making a distributed implementation available, it's currently not in the initial release. Please stay tuned, and take a look at the cifar multi-gpu tutorial for a flavor of how we handle multiple 'devices': http://tensorflow.org/tutorials/deep_cnn/index.md

  3. saraswat commented on Nov 9, 2015

    @saraswat

    would appreciate any insight on the availability of the distributed version. Is the distributed code that is being worked on in github? That is one place where some of us who are interested can contribute

  4. zh4ngx commented on Nov 10, 2015

    @zh4ngx

    👍

  5. edwardyoon commented on Nov 10, 2015

    @edwardyoon

    Hello,

    After reading these plans and ideas, I'm somewhat surprised. According to http://static.googleusercontent.com/media/research.google.com/en//people/jeff/BayLearn2015.pdf, both data and model parallel are needed to train large and powerful models quickly. BTW, GPUs transferring data takes time as described in http://tensorflow.org/tutorials/deep_cnn/index.md. Then, how it's possible to efficiently support both model parallelism and heterogeneous multi-devices (of a single node) on distributed cluster? Could you please roughly explain how different it from DistBelief?

    Thanks!

  6. edwardyoon commented on Nov 10, 2015

    @edwardyoon

    P.S., GPU acceleration also could be limited by model partition strategies.

  7. jeffreyadean commented on Nov 11, 2015

    @jeffreyadean
    Contributor

    Our current internal distributed extensions are somewhat entangled with Google internal infrastructure, which is why we released the single-machine version first. The code is not yet in GitHub, because it has dependencies on other parts of the Google code base at the moment, most of which have been trimmed, but there are some remaining ones.

    We realize that distributed support is really important, and it's one of the top features we're prioritizing at the moment.

  8. edwardyoon commented on Nov 11, 2015

    @edwardyoon

    Awesome.

    After reading the whitepaper, I just realized that large neural network model can be partitioned into sub-graphs by layer (horizontal partitioning) and executed in a serial way.

    One thing not clear is the performance for fully connected network on multi node equipped with GPUs cluster ..

  9. kdunn926 commented on Nov 14, 2015

    @kdunn926

    In theory, something like Dask could be layered on top for handling this - at least for the Python front-end.

  10. saraswat commented on Nov 24, 2015

    @saraswat

    Any update on timeline?

  11. edwardyoon commented on Dec 2, 2015

    @edwardyoon

    Dask looks interesting project but the drawback of the blocking algorithm is that it's not memory optimal. Since a large amount of memory is required for fully-connected layers, I was thought that Pregel-like model parallelism on CPUs w/ vertical partitioning is more attractive for fully connected layers (blocking mat-mult on GPU also appears to me slow and memory demanding). Of course, I maybe wrong but that's why I launched Apache Horn project recently. Since layers can be pipelined, I hope we can collaborate each projects in complementary way.

  12. bhack commented on Jan 15, 2016

    @bhack
    Contributor
  13. saudet commented on Jan 16, 2016

    @saudet

    @bhack I wonder what their Java/Scala interface looks like...

  14. 61 remaining items

  15. added a commit that references this issue on Dec 6, 2019
  16. added a commit that references this issue on Feb 1, 2021
  17. added 5 commits that reference this issue on Apr 9, 2025
  18. added 2 commits that reference this issue on Mar 9, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions