OmniTaskonomy: When Does Visual Generation Improve Visual Understanding

1 University of California, Berkeley 2 Duke University 3 Carnegie Mellon University 4 University of Washington 5 Elorian 6 Impossible Research
* Equal contribution. † Equal contribution.
arXivGitHubHugging Face

TL;DR

Jigsaw I2I
Jigsaw I2T
('reorder', [3, 2, 1, 0])
Visual Generation (Image to Image, I2I)
Ability Transfer
Visual Understanding (Image to Text, I2T)
  1. what training curriculum enables visual generation to improve visual understanding?

  2. which visual generation tasks help which understanding tasks?

  3. What explains the success or failure of transfer between these generation and understanding tasks?

Finding 1

An initial I2I training stage that updates parameters shared with the I2T objective provides a useful initialization for subsequent I2T learning.

Finding 2

Visual generation supervision yields significant gains for specific understanding capabilities, both within and across task families. For example, localization and object pointing yield the largest counting gains, while depth and surface-normal prediction improve metric 3D relation.

Finding 3

Gradient alignment is concentrated in early pre-attention normalization layers and is positively associated with downstream transfer across both understanding capabilities and individual source-target pairs.

When and how does visual generation improve visual understanding?
blue: positivered: negative

Does visual generation help visual understanding?

To isolate the effect of confounding factors, we construct paired I2I and I2T tasks that require solving the same task from the same input, differing only in whether the answer is produced as an image or as text.

Controlled Jigsaw and Zoom-In tasks with image and text outputs.
Controlled Jigsaw and Zoom-In tasks with image and text outputs.

Which training recipe transfers best?

I2T accuracy versus I2I training examples across recipes, with 1k I2T examples. I2T accuracy versus I2T training examples for I2T-only and I2I → I2T training (100k I2I examples). Curves show means over three seeds. Shading indicates ±1 SEM.
I2T accuracy versus I2I training examples across recipes, with I2T examples. I2T accuracy versus I2T training examples for I2T-only and I2I → I2T training ( I2I examples). Curves show means over three seeds. Shading indicates SEM.

I2I → I2T and I2I → Mixed scale consistently as the I2I data increases. Both recipes begin with an I2I-only training stage that updates parameters shared with I2T. In contrast, freezing these weights during the I2I training stage or directly mixing the two objectives from the beginning leads to weaker or less stable gains.

We therefore use I2I → I2T as the default recipe in subsequent experiments.

OmniTaskonomy: A unified taxonomy of visual capabilities

We introduce OmniTaskonomy, a unified taxonomy of generation tasks and understanding capabilities. For a visual understanding sample, we consider the primary visual capability required to solve it; for an I2I task, we consider the visual capability directly supervised by its training objective.

Click a node to explore
OmniTaskonomy
A unified taxonomy of I2I tasks and understanding capabilities under the three Rs: Recognition, Reconstruction, and Reorganization. I2I tasks and understanding capabilities occupy separate leaves in a shared hierarchy.

Which generation tasks help which capabilities?

Hover to magnify · click to pinSwipe to explore · tap a cellNegativePositiveOutlined: (two-sided paired permutation test vs. I2T-only)
I2I supervision task
RecognitionReconstructionReorganization
I2T capabilityObject editingAttribute editingColorizationZ-depthEuclidean depthSurface normalsPrincipal curvatureOcclusion edges3D keypointsReshadingInpaintingSemantic segmentation2D edges2D keypoints2D segmentation2.5D segmentationObject pointingJigsawLocalization
Category recognition
Appearance understanding
Visual similarity
State recognition
Activity understanding
Anomaly detection
Situation understanding
OCR and text recognition
Lighting understanding
Depth understanding
Metric 3D relation
Orientation understanding
2D spatial relation
Multi-view reasoning
Localization
Connectivity
Counting
2D Ordering
Visual correspondence

Visual generation supervision yields significant gains for specific understanding capabilities, both within and across task families. For example, localization and object pointing yield the largest counting gains, while depth and surface-normal prediction improve metric 3D relation.

Shared visual operations predict several of the strongest gains.

Localization and object pointing produce the largest improvements in counting ( and pp), consistent with all three tasks requiring individual objects to be identified and spatially localized.

Similarly, Z-depth, Euclidean depth, and surface normals improve metric 3D relation by , , and pp, respectively, connecting dense geometric reconstruction to relational 3D judgments. Jigsaw improves 2D ordering by pp, consistent with both tasks requiring the relative spatial arrangement of image regions.

Useful transfer is not confined to closely matched capabilities.

Inpainting improves both counting ( pp) and 2D ordering ( pp), despite neither target explicitly requiring missing-region reconstruction. Solving inpainting tasks may encourage the model to infer object quantity and global spatial structure from incomplete local evidence. 2.5D segmentation likewise improves category recognition ( pp), potentially because capturing object boundaries can shape cues that support category recognition.

What explains visual generation-to-understanding transfer?

For each task, we sample matched examples and compute I2I and I2T gradients at the pretrained checkpoint. Because the two objectives solve the same underlying visual problem from the same input, this setting lets us localize where their optimization signals agree.

Module groups
JigsawZoom-In
-0.10.00.20.40.6ViTConn. inConn. out2D pos.Pre-attnQKVQ normK normOutPre-MLPGateUpDown
Pre-attn RMSNormJigsaw 0.54Zoom-In 0.44
Gradient alignment for the controlled Jigsaw and Zoom-In pairs across model components.
RMSNorm layers
JigsawZoom-In
-0.50.00.51.00123456789101112131415161718192021222324252627
Layer 3Jigsaw 1.00Zoom-In 0.76
Gradient alignment for the controlled Jigsaw and Zoom-In pairs across individual pre-attention RMSNorm layers.

Pre-attention RMSNorm parameters matter

For both tasks, alignment is strongest in the understanding branch's pre-attention RMSNorm parameters.

Earlier transformer layers matter

Examining these parameters layer by layer further shows that the strongest alignment occurs in the earlier transformer layers.

We next examine all source–target pairs individually.

For each capability, we sample examples and compute gradients using minibatches of size .

Hover, tap, or focus a point to inspect itRecognitionReconstructionReorganization
Transfer by capability
-2-1012-0.3-0.1500.15Mean gradient alignment across 19 sourcesMean transfer gain (pp)Category recognition — alignment 0.108, transfer +0.44 ppCategory
Appearance understanding — alignment 0.019, transfer +0.01 ppAppearance
Depth understanding — alignment -0.038, transfer +0.07 ppDepth
Metric 3D relation — alignment -0.005, transfer +1.75 ppMetric 3D
2D spatial relation — alignment -0.001, transfer -0.40 pp2D spatial
Counting — alignment 0.186, transfer +1.32 ppCounting
Visual correspondence — alignment -0.246, transfer -2.25 ppCorrespondence

Mean gradient alignment versus mean transfer for seven understanding capabilities, averaging over I2I sources.

Transfer by task pair pairs ·
-6-303-0.5-0.2500.25Gradient alignmentTransfer gain (pp)Object editing → Category recognition — alignment 0.209, transfer +0.07 pp
Attribute editing → Category recognition — alignment 0.202, transfer +1.18 pp
Colorization (Taskonomy) → Category recognition — alignment 0.051, transfer +0.24 pp
Z-depth → Category recognition — alignment 0.125, transfer +1.04 pp
Euclidean depth → Category recognition — alignment 0.137, transfer +0.90 pp
Surface normals → Category recognition — alignment 0.011, transfer +0.14 pp
Principal curvature → Category recognition — alignment 0.105, transfer +0.03 pp
Occlusion edges → Category recognition — alignment 0.074, transfer -0.55 pp
3D keypoints → Category recognition — alignment 0.247, transfer +0.38 pp
Reshading → Category recognition — alignment 0.230, transfer +0.21 pp
Inpainting → Category recognition — alignment 0.142, transfer +0.94 pp
Semantic segmentation → Category recognition — alignment 0.247, transfer +0.66 pp
2D edges → Category recognition — alignment -0.205, transfer +0.69 pp
2D keypoints → Category recognition — alignment 0.042, transfer -0.10 pp
2D segmentation → Category recognition — alignment -0.002, transfer +0.38 pp
2.5D segmentation → Category recognition — alignment 0.084, transfer +1.25 pp
Object pointing → Category recognition — alignment 0.113, transfer -0.14 pp
Jigsaw → Category recognition — alignment 0.107, transfer +0.76 pp
Localization → Category recognition — alignment 0.131, transfer +0.35 pp
Object editing → Appearance understanding — alignment 0.053, transfer -0.82 pp
Attribute editing → Appearance understanding — alignment 0.064, transfer -0.53 pp
Colorization (Taskonomy) → Appearance understanding — alignment -0.038, transfer -0.24 pp
Z-depth → Appearance understanding — alignment -0.120, transfer -0.24 pp
Euclidean depth → Appearance understanding — alignment -0.120, transfer -0.15 pp
Surface normals → Appearance understanding — alignment 0.040, transfer +0.00 pp
Principal curvature → Appearance understanding — alignment 0.136, transfer +0.19 pp
Occlusion edges → Appearance understanding — alignment -0.009, transfer +0.29 pp
3D keypoints → Appearance understanding — alignment 0.222, transfer -0.73 pp
Reshading → Appearance understanding — alignment 0.115, transfer +0.10 pp
Inpainting → Appearance understanding — alignment 0.025, transfer +0.15 pp
Semantic segmentation → Appearance understanding — alignment 0.114, transfer +0.34 pp
2D edges → Appearance understanding — alignment -0.191, transfer +0.44 pp
2D keypoints → Appearance understanding — alignment 0.097, transfer -0.24 pp
2D segmentation → Appearance understanding — alignment -0.048, transfer -0.19 pp
2.5D segmentation → Appearance understanding — alignment 0.059, transfer +0.34 pp
Object pointing → Appearance understanding — alignment -0.025, transfer -0.29 pp
Jigsaw → Appearance understanding — alignment -0.023, transfer +1.50 pp
Localization → Appearance understanding — alignment 0.008, transfer +0.19 pp
Object editing → Depth understanding — alignment 0.014, transfer -0.04 pp
Attribute editing → Depth understanding — alignment -0.005, transfer -0.61 pp
Colorization (Taskonomy) → Depth understanding — alignment -0.073, transfer +0.00 pp
Z-depth → Depth understanding — alignment 0.070, transfer +0.00 pp
Euclidean depth → Depth understanding — alignment 0.147, transfer +0.81 pp
Surface normals → Depth understanding — alignment -0.135, transfer -0.24 pp
Principal curvature → Depth understanding — alignment -0.114, transfer +0.93 pp
Occlusion edges → Depth understanding — alignment -0.095, transfer +0.65 pp
3D keypoints → Depth understanding — alignment -0.080, transfer +0.49 pp
Reshading → Depth understanding — alignment 0.032, transfer -0.12 pp
Inpainting → Depth understanding — alignment -0.065, transfer -0.85 pp
Semantic segmentation → Depth understanding — alignment 0.093, transfer -1.18 pp
2D edges → Depth understanding — alignment -0.076, transfer +0.73 pp
2D keypoints → Depth understanding — alignment -0.067, transfer +0.89 pp
2D segmentation → Depth understanding — alignment -0.126, transfer +0.77 pp
2.5D segmentation → Depth understanding — alignment -0.103, transfer -0.28 pp
Object pointing → Depth understanding — alignment -0.038, transfer -1.99 pp
Jigsaw → Depth understanding — alignment -0.061, transfer +0.41 pp
Localization → Depth understanding — alignment -0.040, transfer +0.89 pp
Object editing → Metric 3D relation — alignment -0.059, transfer +2.03 pp
Attribute editing → Metric 3D relation — alignment -0.086, transfer +0.60 pp
Colorization (Taskonomy) → Metric 3D relation — alignment -0.283, transfer +0.46 pp
Z-depth → Metric 3D relation — alignment 0.045, transfer +3.60 pp
Euclidean depth → Metric 3D relation — alignment 0.169, transfer +3.83 pp
Surface normals → Metric 3D relation — alignment -0.040, transfer +3.41 pp
Principal curvature → Metric 3D relation — alignment -0.002, transfer +1.34 pp
Occlusion edges → Metric 3D relation — alignment -0.096, transfer +1.84 pp
3D keypoints → Metric 3D relation — alignment -0.061, transfer +1.38 pp
Reshading → Metric 3D relation — alignment 0.094, transfer +1.75 pp
Inpainting → Metric 3D relation — alignment -0.072, transfer +1.01 pp
Semantic segmentation → Metric 3D relation — alignment 0.070, transfer +1.94 pp
2D edges → Metric 3D relation — alignment 0.080, transfer +1.20 pp
2D keypoints → Metric 3D relation — alignment 0.137, transfer +0.97 pp
2D segmentation → Metric 3D relation — alignment 0.032, transfer +1.06 pp
2.5D segmentation → Metric 3D relation — alignment 0.044, transfer +2.17 pp
Object pointing → Metric 3D relation — alignment -0.021, transfer +0.32 pp
Jigsaw → Metric 3D relation — alignment -0.019, transfer +2.63 pp
Localization → Metric 3D relation — alignment -0.019, transfer +1.80 pp
Object editing → 2D spatial relation — alignment 0.037, transfer -1.62 pp
Attribute editing → 2D spatial relation — alignment 0.045, transfer -0.58 pp
Colorization (Taskonomy) → 2D spatial relation — alignment -0.176, transfer -1.22 pp
Z-depth → 2D spatial relation — alignment -0.137, transfer -0.79 pp
Euclidean depth → 2D spatial relation — alignment -0.102, transfer -0.58 pp
Surface normals → 2D spatial relation — alignment 0.008, transfer -0.58 pp
Principal curvature → 2D spatial relation — alignment 0.079, transfer +0.40 pp
Occlusion edges → 2D spatial relation — alignment -0.078, transfer -1.11 pp
3D keypoints → 2D spatial relation — alignment 0.164, transfer +0.07 pp
Reshading → 2D spatial relation — alignment 0.068, transfer +0.00 pp
Inpainting → 2D spatial relation — alignment 0.035, transfer -0.40 pp
Semantic segmentation → 2D spatial relation — alignment 0.108, transfer -0.32 pp
2D edges → 2D spatial relation — alignment -0.170, transfer +0.50 pp
2D keypoints → 2D spatial relation — alignment 0.140, transfer +0.32 pp
2D segmentation → 2D spatial relation — alignment -0.060, transfer -0.07 pp
2.5D segmentation → 2D spatial relation — alignment 0.006, transfer -0.68 pp
Object pointing → 2D spatial relation — alignment -0.008, transfer -0.14 pp
Jigsaw → 2D spatial relation — alignment -0.005, transfer -0.22 pp
Localization → 2D spatial relation — alignment 0.024, transfer -0.58 pp
Object editing → Counting — alignment 0.282, transfer +1.08 pp
Attribute editing → Counting — alignment 0.270, transfer +1.45 pp
Colorization (Taskonomy) → Counting — alignment 0.085, transfer +0.73 pp
Z-depth → Counting — alignment 0.259, transfer +1.58 pp
Euclidean depth → Counting — alignment 0.245, transfer +0.97 pp
Surface normals → Counting — alignment 0.086, transfer +1.28 pp
Principal curvature → Counting — alignment 0.195, transfer +0.82 pp
Occlusion edges → Counting — alignment 0.140, transfer +1.02 pp
3D keypoints → Counting — alignment 0.311, transfer +1.49 pp
Reshading → Counting — alignment 0.303, transfer +1.19 pp
Inpainting → Counting — alignment 0.200, transfer +1.49 pp
Semantic segmentation → Counting — alignment 0.317, transfer +1.08 pp
2D edges → Counting — alignment -0.161, transfer +1.53 pp
2D keypoints → Counting — alignment 0.074, transfer +1.41 pp
2D segmentation → Counting — alignment 0.160, transfer +1.43 pp
2.5D segmentation → Counting — alignment 0.206, transfer +1.06 pp
Object pointing → Counting — alignment 0.191, transfer +1.97 pp
Jigsaw → Counting — alignment 0.173, transfer +0.95 pp
Localization → Counting — alignment 0.198, transfer +2.49 pp
Object editing → Visual correspondence — alignment -0.319, transfer -2.03 pp
Attribute editing → Visual correspondence — alignment -0.287, transfer -2.30 pp
Colorization (Taskonomy) → Visual correspondence — alignment -0.052, transfer -0.92 pp
Z-depth → Visual correspondence — alignment -0.520, transfer -1.38 pp
Euclidean depth → Visual correspondence — alignment -0.501, transfer -4.59 pp
Surface normals → Visual correspondence — alignment -0.026, transfer -2.89 pp
Principal curvature → Visual correspondence — alignment 0.039, transfer -1.31 pp
Occlusion edges → Visual correspondence — alignment -0.173, transfer -0.59 pp
3D keypoints → Visual correspondence — alignment -0.207, transfer -4.33 pp
Reshading → Visual correspondence — alignment -0.289, transfer -3.02 pp
Inpainting → Visual correspondence — alignment -0.273, transfer -1.84 pp
Semantic segmentation → Visual correspondence — alignment -0.186, transfer -4.00 pp
2D edges → Visual correspondence — alignment -0.244, transfer -2.62 pp
2D keypoints → Visual correspondence — alignment -0.155, transfer -0.26 pp
2D segmentation → Visual correspondence — alignment -0.273, transfer -3.61 pp
2.5D segmentation → Visual correspondence — alignment -0.256, transfer -5.51 pp
Object pointing → Visual correspondence — alignment -0.333, transfer +1.71 pp
Jigsaw → Visual correspondence — alignment -0.318, transfer -2.89 pp
Localization → Visual correspondence — alignment -0.301, transfer -0.33 pp
2D edges → Category recognitionObject pointing → Counting3D keypoints → Appearance2.5D segmentation → Correspondence

Alignment and transfer for all source–target pairs.

Alignment predicts capability gains

Average alignment and average transfer are strongly positively correlated across the seven capabilities (). Understanding capabilities whose gradients are more compatible with generation objectives also tend to receive larger gains from I2I training.

Does alignment track transfer across task pairs?

Alignment and transfer are positively correlated across these pairs (). Together, these results identify gradient alignment as an optimization-level signal associated with generation-to-understanding transfer.

Citation

@article{ge2026omnitaskonomy,
  title={OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?},
  author={Ge, Jiaxin and Qin, Yiming and Xie, Ji and Jiang, Haozhe and Han, Xiaochuang and Zhang, Junyi and Dai, Andrew and Yang, Yinfei and Malik, Jitendra and Krishna, Ranjay and Min, Sewon and Feng, Haiwen and Xue, Le and Shi, Baifeng and Darrell, Trevor and Wang, XuDong},
  journal={arXiv preprint arXiv:2609.38079},
  year={2026}
}