Latent Reasoning

Research on language models that reason in continuous space instead of words.

A language model that reasons step by step writes each step down. The step is rounded to the nearest token, emitted, and read back in. That rounding is not incidental. It is a quantiser sitting in the middle of the reasoning loop, and it throws away almost everything the model computed. Latent reasoning is the study of what happens when you take it out.

The question has two halves. Removing the quantiser widens the channel between reasoning steps by three orders of magnitude, which buys parallel search, shorter traces, and computation that no vocabulary can express. It also removes the bound that currently forces a model to route serial reasoning through text, which is not the same as saying the text was ever a faithful account of the reasoning, and the evidence that it was not is substantial.

Token space Chain of thought round each step to a word Latent space Continuous thought feed the hidden state back the quantiser the only difference β ≈ 17 bits one token per step β ≈ 10⁴ bits one hidden state per step β in between largely unexplored written down not written down

One axis: how many bits of each reasoning step survive to the transcript. Whether what survives can be read is a separate question the axis does not settle, taken up in the introduction.

A transcript that determines a computation is a long way from a transcript that explains it. The introduction works through the mechanism and its limits.

The state of play, September 2026

Twenty-one months separate the naming of continuous thought in December 2024 from a production controversy, with four dedicated surveys arriving along the way. What the literature has settled depends on which claim is being made.

Efficiency, split four ways

Depth-recurrent and latent methods beat token-matched explicit chain of thought on reported accuracy per inference FLOP and on thought-phase latency, and open weights exist at 1.4B to 3.5B parameters. But parameter savings, training compute, inference compute and wall-clock latency are four different ledgers, and a win in one is not a win in the others. The iso-depth scaling study prices a recurrence at roughly $r^{0.46}$ equivalent unique parameters and finds a training-compute cost to parameter sharing in its setup. Claims here should say which ledger they are drawn on.

The mechanism claims are contested

The theoretical selling point of continuous thoughts is superposition, holding several reasoning paths at once rather than committing to one. That mechanism is proven for a specific construction and demonstrated in models trained from scratch. It was also shown not to occur in one prominent method that claimed it, and shown to collapse under the fine-tuning regimes practitioners actually use. Recurrent depth does not scale monotonically either. Looped performance can peak and then fall with further iterations (Yang et al. 2026), and extra steps can move in the right direction by the wrong distance (Guo et al. 2026). Huginn improves monotonically out to 64 iterations, so this describes some looped models rather than recurrence as such.

Continuous and discrete steps separate, conditionally

Continuous latent steps reach $\mathsf{TC}^k$ where discrete steps at matched depth reach only $\mathsf{TC}^{k-1}$: under deterministic decoding. Under stochastic decoding, discrete chains win back counting and sampling power. The two are reported as complementary resources rather than a ranking. The theory page has the statements and their premises.

Hybrids are the thinnest part of the literature

Almost all work sits at one endpoint or the other. Methods that interleave latent and textual steps are where the efficiency and oversight arguments would meet directly, and there are few of them.

Deployment

In September 2026 the question stopped being architectural. OpenAI released GPT-6 Astra and its system card records “a substantial decrease in chain-of-thought monitorability compared to previous models”: a frontier lab documenting the loss as a measured property of a shipped model. The card does not name an architecture. The attribution to recurrent depth comes from reporting, and the two claims should be kept apart. The position paper that anticipated all of this was published thirteen months earlier by forty authors drawn from the labs now on both sides of it.

That argument cuts in both directions, and the monitorability page sets it out with the measurements attached: chains of thought were never as faithful as the safety case implies, and losing an unreliable channel before its replacement is ready is still a loss.

Further reading, by depth

The annotated bibliography lives on its own page, across theory, continuous thoughts, recurrent depth, contentless tokens, rebuttals and monitorability, each entry with a note on why it matters. The bibliography has them all, and reviews gives a handful a closer reading.