Repository navigation
API docs does not list RNNs #7
Description
Activity
You should look at the RNN tutorial: http://tensorflow.org/tutorials/recurrent/index.md .
(Thanks, I have seen the tutorial but it is not a substitute for an API; I filed the issue after consulting with a friend at Google Brain)
Hi Avanti -- internally we've been working on iterating the API for RNNs, and we were happy enough with the current API to use it in the tutorial, but we're making sure it's solid before promoting it to the public API, since we'd then have to support it indefinitely. (Anything not in the public API is a work-in-progress :)
We'll keep this bug open in the meantime, and for now you can look at the source code documentation if you're interested in playing around: https://github.com/tensorflow/tensorflow/blob/master/tensorflow/models/rnn/rnn.py#L9
Ah, got it, thanks for the explanation.
The white paper of TensorFlow mentions looping control within the graph. Is it already available? If so, are there examples to show how it can be done?
The RNN example has a Python loop. Will TensorFlow treat that as a symbolic loop and compile it?
Also, the explanation of
sequence_lengthhere isn't clear to me. What does it mean by dynamic calculations? Whentis pastmax_sequence_length, can it just break from the loop instead of continuing withzerosstate? Returningzerosstate is different from returning the state atmax_sequence_length, isn't it?On your first question, see #208.
On your second question: the core TF engine currently only sees the GraphDef produced by python, so the RNN example is an unrolled one today.
I'm not super familiar with that RNN example -- @lukaszkaiser or @ludimagister might know better.
I zer0n,
the current RNN is statically unrolled, there is no (not yet) dynamic unrolling based on the length of the sequence. Th dynamic calculation means the graph is unrolled up to the max_sequence_length, but if a sequence_length is provided the calculations on the unrolled graph are cut short once the sequence_length is reached, using a conditional op. Depending on the application this may result in shorter processing time.
Yes, to add to what @ludimagister says: the conditional op will plug in zeros to the output & state past max(sequence_length), thus reducing the total amount of computation (if not memory).
I may actually modify this so that instead of plugging in zeros to the state, it just copies the state. This way the final state will represent the "final" state at max(sequence_length). However, I'm undecided on this. If you want the final state at time sequence_length, you can concat the state vectors and use transpose() followed by gather() with sequence_length in order to pull out the states you care about. That's probably what you would want to do, in fact, because if you have batch_size = 2 and sequence_length = [1, 2], then for the first minibatch entry, the state at max(sequence_length) will not equal the state at sequence_length[0].
An alternative solution is to right-align your inputs so that they always "end" on the final time step. This breaks down the dynamic calculation performed when you pass sequence_length (because it assumes left-aligned inputs). I may extend this by adding a bool flag like "right_aligned" to the rnn call, which assumes that calculation starts at len(inputs) - max(sequence_length), and copies the initial state through appropriately. But that doesn't exist now.
Thanks @vrv, @ludimagister, and @ebrevdo for the answers. However, some details still confuse me.
- @ludimagister, the code doesn't seem to statically unroll. It has a loop which depends on the length of the inputs. Plus,
max_sequence_lengthis not a const; instead it's just the scalar of thesequence_lengthparameter, which can be and isNoneby default. So, by default, the unrolling is not truncated. Correct me if I misread the code. - @ebrevdo I understand the computational saving motivation. However, returning zeros is logically very different from returning the state at
sequence_length(if provided). The former is just wrong. Again, please correct if I misread the code.
Thanks.
- @ludimagister, the code doesn't seem to statically unroll. It has a loop which depends on the length of the inputs. Plus,
@zer0n It depends on your task.
Returning zeros is fine if you only care about outputs (i.e., you're not hooking up to a decoder); and your loss function knows to ignore outputs past the sequence_length.
Returning the state from the end of the last time step might also be considered "wrong", but will generally always happen if you have inputs of different lengths (and aren't performing dynamic computation). This is a typical approach to performing RNN with minibatches. For this reason when performing encoding/decoding, people usually right-align with left-side padding instead, so the last input of any example always corresponds to the very last state. This seems like the cleanest solution for now.
Anyway, this part of the API may change; not sure yet the best approach.
(also, specifically returning the state at sequence_length for every entry is taxing both in terms of computation and in terms of memory, both in short supply with RNNs )
OK, I did miss the line
outputs.append(output). I originally thought that it returned the final state, not a sequence of states.Anyway, this implementation still looks weird (I'm aware it's changing so I'm only discussing the current state). Usually, for truncated BPTT implementation, people pad
eosfor short sentences and truncate the sentences if the lengths are larger thanmax_length. This enables static unrolling and efficient mini-batching.The RNN example seems doing the reverse. What I see is that it's doing dynamic unrolling (i.e. with dynamic output size), but padding zeros to the outputs past
max_length.(This discussion is probably better off had on the discussion mailing list, rather than this bug about documentation)
@vrv:
but we're making sure it's solid before promoting it to the public API, since we'd then have to support it indefinitely.
May we assume that TF is going to use semantic versioning for releases (ie. major.minor.patch)?
Major releases can have backward-incompatible API changes, and minor releases can certainly add a new API (or extend an existing one in a backward-compatible way) especialy if it marks the old API as deprecated (and to be removed in the next major release).
Since TF has not yet had a major version release, (the current release is only 0.5.0), you have a lot of wiggle room between now and an eventual 1.0.0 release that then really would commit you to maintaining backward compatibility for quite a while.
50 remaining items
- added 7 commits that reference this issue
on Apr 9, 2025 - added 5 commits that reference this issue
on Jul 28, 2025 - added a commit that references this issue
on Nov 26, 2025
https://github.com/tensorflow/tensorflow/blob/master/tensorflow/g3doc/api_docs/python/nn.md
(It has convolutional layers listed, for instance, but does not show the RNNs. Actually, I don't quite see anything corresponding to a fully-connected layer either)