Mean gradient alignment versus mean transfer for seven understanding capabilities, averaging over I2I sources.
Does visual generation help visual understanding?
To isolate the effect of confounding factors, we construct paired I2I and I2T tasks that require solving the same task from the same input, differing only in whether the answer is produced as an image or as text.
Which training recipe transfers best?
I2I → I2T and I2I → Mixed scale consistently as the I2I data increases. Both recipes begin with an I2I-only training stage that updates parameters shared with I2T. In contrast, freezing these weights during the I2I training stage or directly mixing the two objectives from the beginning leads to weaker or less stable gains.
We therefore use I2I → I2T as the default recipe in subsequent experiments.
OmniTaskonomy: A unified taxonomy of visual capabilities
We introduce OmniTaskonomy, a unified taxonomy of generation tasks and understanding capabilities. For a visual understanding sample, we consider the primary visual capability required to solve it; for an I2I task, we consider the visual capability directly supervised by its training objective.
Which generation tasks help which capabilities?
| I2I supervision task | |||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Recognition | Reconstruction | Reorganization | |||||||||||||||||
| I2T capability | Object editing | Attribute editing | Colorization | Z-depth | Euclidean depth | Surface normals | Principal curvature | Occlusion edges | 3D keypoints | Reshading | Inpainting | Semantic segmentation | 2D edges | 2D keypoints | 2D segmentation | 2.5D segmentation | Object pointing | Jigsaw | Localization |
| Category recognition | |||||||||||||||||||
| Appearance understanding | |||||||||||||||||||
| Visual similarity | |||||||||||||||||||
| State recognition | |||||||||||||||||||
| Activity understanding | |||||||||||||||||||
| Anomaly detection | |||||||||||||||||||
| Situation understanding | |||||||||||||||||||
| OCR and text recognition | |||||||||||||||||||
| Lighting understanding | |||||||||||||||||||
| Depth understanding | |||||||||||||||||||
| Metric 3D relation | |||||||||||||||||||
| Orientation understanding | |||||||||||||||||||
| 2D spatial relation | |||||||||||||||||||
| Multi-view reasoning | |||||||||||||||||||
| Localization | |||||||||||||||||||
| Connectivity | |||||||||||||||||||
| Counting | |||||||||||||||||||
| 2D Ordering | |||||||||||||||||||
| Visual correspondence | |||||||||||||||||||
Visual generation supervision yields significant gains for specific understanding capabilities, both within and across task families. For example, localization and object pointing yield the largest counting gains, while depth and surface-normal prediction improve metric 3D relation.
Shared visual operations predict several of the strongest gains.
Localization and object pointing produce the largest improvements in counting ( and pp), consistent with all three tasks requiring individual objects to be identified and spatially localized.
Similarly, Z-depth, Euclidean depth, and surface normals improve metric 3D relation by , , and pp, respectively, connecting dense geometric reconstruction to relational 3D judgments. Jigsaw improves 2D ordering by pp, consistent with both tasks requiring the relative spatial arrangement of image regions.
Useful transfer is not confined to closely matched capabilities.
Inpainting improves both counting ( pp) and 2D ordering ( pp), despite neither target explicitly requiring missing-region reconstruction. Solving inpainting tasks may encourage the model to infer object quantity and global spatial structure from incomplete local evidence. 2.5D segmentation likewise improves category recognition ( pp), potentially because capturing object boundaries can shape cues that support category recognition.
What explains visual generation-to-understanding transfer?
For each task, we sample matched examples and compute I2I and I2T gradients at the pretrained checkpoint. Because the two objectives solve the same underlying visual problem from the same input, this setting lets us localize where their optimization signals agree.
Pre-attention RMSNorm parameters matter
For both tasks, alignment is strongest in the understanding branch's pre-attention RMSNorm parameters.
Earlier transformer layers matter
Examining these parameters layer by layer further shows that the strongest alignment occurs in the earlier transformer layers.
We next examine all source–target pairs individually.
For each capability, we sample examples and compute gradients using minibatches of size .
Alignment and transfer for all source–target pairs.
Alignment predicts capability gains
Average alignment and average transfer are strongly positively correlated across the seven capabilities (). Understanding capabilities whose gradients are more compatible with generation objectives also tend to receive larger gains from I2I training.
Does alignment track transfer across task pairs?
Alignment and transfer are positively correlated across these pairs (). Together, these results identify gradient alignment as an optimization-level signal associated with generation-to-understanding transfer.
Citation
@article{ge2026omnitaskonomy,
title={OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?},
author={Ge, Jiaxin and Qin, Yiming and Xie, Ji and Jiang, Haozhe and Han, Xiaochuang and Zhang, Junyi and Dai, Andrew and Yang, Yinfei and Malik, Jitendra and Krishna, Ranjay and Min, Sewon and Feng, Haiwen and Xue, Le and Shi, Baifeng and Darrell, Trevor and Wang, XuDong},
journal={arXiv preprint arXiv:2609.38079},
year={2026}
}

