Preprint · 2026

RelateAnything

Any region·any relation·in real time

An image, a set of boxes, and a list of relation names as text go in. A score for every pair of boxes against every name comes out, at 20 ms per frame. The model is never given object class labels, and it runs in your browser.

Maëlic Neau · Independent Researcher

Weights on Hugging Face · the corpus as RA-4M

real model output

Each graph is the 53M model's own output on a Creative Commons photo. Run it on your own images; nothing is uploaded.

One 45-second clip, three ways of drawing the regions: boxes, then instance masks, then those masks with the names hidden. The same ONNX graph as the demo throughout, handed box coordinates and pixels, never words — and a filter on calibrated log-odds holds each relation through the frames where an endpoint goes undetected.

19,103predicate strings answerable at inference; VG150 uses 50
53Mparameters, 20 ms per frame on an A40 with the detector
0object class labels, at any stage of training or inference
474kimages and 4.3M verified relations, generated in 104 GPU-hours

Detection and segmentation have already moved the taxonomy out of the model and into the input. Relation prediction has not: scene-graph models still learn one corpus's 50 predicates, conditioned on object class labels. RelateAnything takes pixels, regions from any source, and a vocabulary supplied as strings. On three benchmarks that contributed no training image and a fourth evaluated zero-shot, it has 2.3–3.5× the mean recall of the strongest open-vocabulary method of comparable scale and 5–21× its rare-predicate recall, at 7.8× the speed; the six-axis composite is 40.1 against 11.8. It also leads a scene-graph model built on a 3B vision–language model on both recall metrics on all three benchmarks both were run on, at under 2% of its parameters.

Getting there meant measuring the metric first. A frequency table that never sees an image beats this model on the number leaderboards are ranked by — so the benchmark here reports six axes, two of them scored against negatives a human adjudicated.

Part I · Corpus

RA-4M 474k images · 4.3M relations

An open-weight VLM annotates images marked with a numbered dot per box, in three passes; deterministic geometric checks then reject 11.3% of what it writes. 10,102 free-text predicates, and 9.03 relations per image against the source annotations' 5.29 on the same images and boxes.

We know of no other machine-annotated relation corpus that checks its annotator against box geometry.

Part II · Model

RelateAnything DINOv3 ViT-S/16+ · 53M

Pixels and box coordinates in, one calibrated score per predicate string out. The strings are encoded by a distilled text encoder into an ℓ₂-normalised bank that no learned layer touches.

Adding a predicate means adding a row. Regions can come from any detector, or from a segmenter with no class names at all.

Part III · Benchmark

OV-SGG-Bench six axes, one composite

Transfer, precision, open vocabulary, deployment, graph quality, spatial — all scored across datasets. Each can be satisfied on its own by a different shortcut, which is why all six are reported together.

Every recall number is reported beside the shared triplet mass between training corpus and benchmark.

Interactive · nothing to download

The predicate list is an input

Pick a photo and choose which relations to score. The model scored 263 predicate strings on each of these photos once; toggling a word re-runs the demo's decode on those stored scores. Switch the region source to FastSAM and the objects lose their names entirely — the graph is computed the same way.

Objects from

Predicates

● spatial ● semantic. A word with a coloured ring appears in the current graph. To type a predicate that is not listed, use the demo.

Results · cross-dataset, no object labels

40.1 against 11.8 on the composite

Six axes, each scored on benchmarks that contributed no training image, each against the strongest system that can be run on it. Pick an axis; the chart carries the numbers and the notes carry the caveats.

axis
metric
view
RelateAnything, ViT-S/16+ (no object labels)OvSGTR (given ground-truth object labels)
What each axis is, and how each one fails when read alone
  • A1

    Transfer. Closed-vocabulary recall, ground-truth boxes, four sources of differing shared triplet mass.Alone: gameable by corpus match.

  • A2

    Precision. Federated AP against adjudicated negatives, the only axis that sees a confident false positive.Alone: invariant to uniform score depression.

  • A3

    Open vocabulary. Full training vocabulary deployed, synonym-tolerant matching: does the model mean the right relation?Alone: rewards head collapse onto on / in.

  • A4

    Deployment. Detection-mode graphs on a shared detector, against the measured pair-recall ceiling.Alone: bounded by the detector rather than the model.

  • A5

    Graph quality. A VLM judge, one relation at a time, no ground truth in the prompt; each acceptance credited with its surprisal.Alone: shown whole graphs, the same judge prefers short repetitive ones. That verdict is reported too.

  • A6

    Spatial. SpatialSense: adversarial, balanced true/false, where knowing that on is common does not help.Alone: a boxes-only baseline scores 68.8 without the image.

The composite is the chance-corrected harmonic mean over A1, A2, A4, A5 and A6, with A1 entering as F1@50 averaged over its four benchmarks. It spans only the axes a baseline can be run on, so OvSGTR alone carries it; A3 is reported and left out, because the one baseline runnable there is scored through a matcher whose choice moves the result by more than the two systems differ.

Per-predicate spatial AUC across model families
Per-predicate AUC on SpatialSense (A6). on and in are learned. above, to the left of and next to sit at chance, and fifteen months of recipe changes did not move them. This axis remains the weakest.
Three comparisons beyond the six axes: the OvR-SGG leaderboard, the full 19,103-string vocabulary, and systems that answer in free text

On the OvR-SGG leaderboard

OvR-SGG holds 15 of VG150's 50 predicates out of training and scores in detection mode. We reproduce all twelve published numbers to within 0.35 points, then retrained with those 15 strings and their synonym groups removed.

Methodbackbone / boxesB+N R@50Novel R@50B+N mR@50Novel mR@50classes hit
VS3Swin-T15.60.0n/rn/rn/r
OvSGTRSwin-T20.513.5n/rn/rn/r
RAHPSwin-T20.515.6n/rn/rn/r
OvSGTR + MegaSGSwin-T25.417.0n/rn/rn/r
OvSGTRSwin-B22.916.4n/rn/rn/r
INOVASwin-B24.820.0n/rn/rn/r
OvSGTR re-run hereSwin-T20.413.24.11.913/50 3/15 novel
RelateAnything novel 15 held outtheir boxes22.411.814.97.240/50 8/15 novel
RelateAnything zero-shot tower, novel set seentheir boxes27.322.8n/rn/rn/r
RelateAnything released, novel set seentheir boxes27.924.1n/rn/rn/r

On the leaderboard's micro metric the held-out model is mid-table, and we claim no lead there. On mean recall over the same predictions it is nearly four times the baseline, answering with 40 of the 50 predicates against 13. The split explains both: the 15 held-out predicates carry 55% of the test relations, and on, of and in account for 93% of that. The last two rows have those 15 strings in training, so their Novel columns do not satisfy the protocol; the gap to them, −5.5 and −12.3 R@50, measures how much of novel performance here is supervision on those three words. Mean recall and the class count are measured here for the two systems we run and reported by no published row (n/r).

With all 19,103 strings deployed

Exact-string matching penalises an open-vocabulary model for its own synonyms: to the right of scores zero against right of. With synonym-tolerant matching over the full training vocabulary, the correct relation is the model's median first choice on two of the three benchmarks.

SourceR@50mR@50rareMRRmedian rank
VG15056.034.539.60.681 / 19,103
PSG30.528.320.80.3910 / 19,103
IndoorVG53.334.634.60.651 / 19,103

For scale, the published closed-vocabulary ceilings under PredCls are 68.2 R@50 (PE-Net) and 37.0 mR@50 (VCTree+IETrans+Rwt) on VG150, and 41.7 mR@50 on PSG (DSFlash-L, under the corrected protocol of Lorenz et al.) — from specialists trained on that benchmark, with ground-truth object labels and a fixed head.

Against systems that answer in free text

A system that answers in free text can be scored on A3 by construction, which is the one comparison OvSGTR cannot enter: its vocabulary arrives as a single caption, about 150 strings. That admits ROBIN-3B, a scene-graph model built on a 3B vision–language model, and general multimodal models prompted for the same output. Below is one set of generations per model read several ways, so a system's rows differ only in how its free text is mapped onto the benchmark's words.

ModelmatcherR@20R@50mR@20mR@50pairs reached
PSG test, 2,179 images, ground-truth regions for every model, one scorer
RelateAnything 53Mexact string12.313.612.413.099.7
RelateAnythingsynonym, τ = 0.6031.736.829.131.399.7
ROBIN-3Bexact string25.826.219.920.045.6
ROBIN-3BSBERT argmax its own30.435.721.122.977.4
ROBIN-3Bsynonym, τ = 0.6035.137.924.125.165.6
Qwen3-VL-32B promptedsynonym, τ = 0.6015.015.65.15.235.5
Qwen3-VL-8B promptedsynonym, τ = 0.6010.210.23.33.323.6
InternVL3.5-8B promptedsynonym, τ = 0.609.59.64.54.523.1
GLM-4.6V-Flash promptedsynonym, τ = 0.608.99.12.12.223.4

The matcher is a scoring convention whose effect exceeds the difference it is used to measure: mapping free text onto the benchmark's words is worth +11.7 R@50 to ROBIN and +23.2 to us, since a model answering from 19,103 strings is the one exact matching penalises hardest. It also decides the mean-recall ordering — ROBIN leads on exact strings, 20.0 against 13.0, and we lead under every synonym-tolerant matcher, 31.3 against 25.1. Two things survive it. One is pairs reached, the share of annotated pairs a system names at all: 99.7% for us, 45.6–77.4% for ROBIN, 23.1–35.5% for the prompted models, and a pair never named cannot carry the right predicate. The other is what happens off PSG: run on three benchmarks, each under its own protocol, we lead ROBIN on both metrics on all three, by 1.9× on VG150 micro recall and 2.2× on IndoorVG against 1.1× on PSG.

The model

One score per predicate string, from pixels and box coordinates

A visual path and a text path that meet at a cosine. The figure is the paper's own; the chips under it take the blocks one at a time.

The RelateAnything architecture in three bands. Band 1, the visual path: one frame goes both to the region producer, any detector or segmenter, and to a fine-tuned DINOv3 backbone; a patch map fused from three depths is read by one query per box, five pooled features form each ordered pair, and 128 sampled pairs are refined by a relation transformer and a deformable read of the scene. Band 2, the text path: predicate strings are embedded once by a frozen distilled encoder into a bank of unit vectors, a gate reads the mix between the geometry-led and appearance-led branches from the string alone, and the score is a cosine to the predicate embedding, mixed by that gate, plus the pair logit, through a sigmoid. Band 3: the training loss on one annotated pair, and the scored relations that come out.
The model, as the paper draws it. The frame goes to the backbone and, unchanged, to whatever produces the regions; that producer is not part of the model and its class labels are discarded. Predicate strings supplied at inference are embedded once into a bank of unit vectors that no learned layer transforms, and a gate maps each embedding to the weight αp mixing the two branches. Band 3 is training only.

No object labels

The standard protocol supplies ground-truth boxes and their classes, so a model can lean on the prior that a ⟨person, bicycle⟩ pair is probably riding. Ours sees pixels and box coordinates only, which is the situation when boxes come from a live detector. Measured: 87–93% of the semantic logit's variance comes from pair context and 0.1% from object identity, and scoring against another image's features costs 44–68% of micro accuracy.

Ten thousand predicates change the loss

At this vocabulary size the supervision is positive-unlabeled: a ⟨man, horse⟩ pair annotated riding is also sitting on, and scoring that as a negative teaches the model to suppress correct answers. Each negative is therefore discounted by the estimated probability that it also holds — fitted on the 706k training pairs carrying more than one annotation — while annotated predicates and directional inverses keep full weight, since separating above from below is the one signal that must not be softened.

The text side is never learned

A distilled 512-d encoder maps each string into the head's space, and nothing there is trained after the bank is built, so a string never seen in training is treated exactly like one that was. It has to be distilled because contrastive encoders are direction-blind: the teacher puts above and below at cosine 0.95 and the two left/right forms at 0.99, indistinguishable from its synonym pairs at 0.96. No visual model regressing onto such targets can separate them, whoever trains it.

The corpus

RA-4M: machine annotation, deterministically verified

Human relation annotation is expensive and annotators agree with each other poorly. RA-4M generates annotations at scale and then filters them against box geometry.

RA-4M generation pipeline: three VLM passes over images marked with numbered dots, then deterministic geometric gates, then a salience-ranked top-up
Generation pipeline. An open-weight 26B mixture-of-experts VLM (4B active) annotates images carrying a numbered dot at the centre of each box, in three passes. Geometric gates then reject 11.3% of the raw candidates. 104 GPU-hours end to end.

A gate fires only where box geometry logically constrains the predicate, so a rejection is a true negative up to box error, and what geometry cannot constrain passes through unchecked and is counted rather than guessed at. Contact predicates must touch; containment and proximity are checked against the boxes; direction is repaired rather than rejected, the roles swapped to match the geometry with the annotator's own wording kept. Because the annotator prefers the left and upper object as subject, each verified directional relation is then restated from the other endpoint with probability one half — roles swapped, predicate replaced by its inverse — which leaves the corpus direction-balanced, a property the model depends on and no source corpus has.

Corpusimagesrelationsrel/imgpredicatesentropy
RA-4M474,4134,282,5319.0310,1023.99
MegaSG source annotations, same images474,4202,510,9055.29942.56
raw Visual Genome108,0772,315,90622.1036,5493.81
  after the leakage filter (secondary source)40,615758,54718.7017,742—

Against the annotations it replaces, on the same images and the same boxes: 1.7× denser, 107× the vocabulary, 1.4 nats more predicate entropy, and 73 of the 94 source predicate classes reproduced as exact strings. Mean object degree rises from 1.88 to 3.20 and the share of images whose graph is a single connected component from 70.9% to 82.6%. The largest unverifiable class is gaze: about 22% of looking at and watching annotations have disjoint boxes.

What the vocabulary looks like, and the same images under both annotators
Top-50 predicates
Top-50 predicates. Free text, surface forms kept. behind, in front of, wearing and the two left/right forms, balanced between inverses by construction, replace the source's near and on.
Predicate frequency, Zipf plot
The tail. Rank–frequency of all 10,102 predicates. The top ten still carry 57% of the mass, but ranks 10–100 hold an order of magnitude more than the source vocabulary does.
MegaSG annotations versus RA-4M on the same images
Same images, two annotators. The released MegaSG annotations (94 predicates) beside RA-4M. Every evaluation split is perceptually de-duplicated against the corpus, under both of the id spellings Visual Genome images are distributed with.

Measurement

What scene-graph recall measures

Four results that changed how we read every number above.

A frequency table beats our model on the ranked metric

The freq baseline looks up the most frequent training predicate for a ⟨subject, object⟩ category pair. It reads the ground-truth object labels and never sees the image. It wins on per-edge accuracy on all three benchmarks — 68.4 against 57.7 on VG150 — and loses on the per-predicate average of the same predictions, 18.9 against 35.1. A metric a lookup table can win does not measure relation understanding, and it is the metric that orders leaderboards.

Top-1 over ground-truth pairs and boxes, the model being the tower trained without any HICO-DET share. 6.9–12.8% of edges are ones we get right and it gets wrong, so pixels are clearly doing something.

What a benchmark shares with a training corpus is not the vocabulary but the triples

on between a person and a horse and on between a book and a table are different acts; two corpora can agree on the string and never agree on the pair it is asserted of. Shared triplet mass is the share of a training corpus's relation instances whose ⟨subject category, predicate, object category⟩ triple the benchmark also annotates. It collapses when the object categories are charged for, and it collapses unevenly: on VG150 the baseline's fine-tuning corpus scores 90.9% by that measure, against our 12.8%. The confound is concentrated in-domain, and it is very large there.

The full decomposition, ours against the baseline's fine-tuning corpus
Share of the corpus's relations matched onVG150PSGIndoorVGHaystack
our released training mixture · 19,103 predicate strings
the predicate string53.438.749.538.7
both object categories44.436.726.830.7
the whole triple12.810.76.64.9
VG150 train · 50 predicates · the baseline's fine-tuning corpus
the predicate string100.057.395.757.3
both object categories100.022.112.37.1
the whole triple90.98.610.40.6

No image and no annotation is shared with any benchmark; leakage is excluded separately, by image identity. The statistic is predictive — one arm of ours with a larger share of raw Visual Genome reached 54.3 R@50 on VG150, the best zero-shot figure we know of, and was also the worst model we trained on every tail metric.

Two scoring conventions move results more than the methods do

The matcher grants duplicate credit.

The standard detection-mode matcher has no assignment constraint, so a ground-truth object covered by d detections gives d chances at the same relation. On VG150 test the mean is 2.99. Adding the constraint costs the published baseline 2.45 R@50 points; upgrading its backbone from Swin-T to Swin-B is worth 2.39.

The detector's operating point is unreported.

Detection-mode recall is bounded by pair recall, which is quadratic in object recall. With the detector weights fixed, moving only the confidence threshold and the box cap shifts that ceiling by 19 points on VG150. The spread between all published methods on that leaderboard is 4.3 points.

One cheap fix covers most of this: state the matcher, the detector's threshold and box cap, and the pair-recall ceiling they imply, next to every detection-mode number.

Measuring in-domain overstates transfer by about 5×

Every ablation was read twice, once on the corpus's own validation split and once on benchmarks the model never trained on. Doubling the training corpus is worth +15% micro and +56% macro recall in-domain, and +2.5–3.8% and +9.9–11.9% under transfer: the extra images teach the corpus's annotation style, which the in-domain split shares by construction. One arm changes sign outright — removing an automatically derived left/right excess improves in-domain metrics at every epoch and lowers transfer macro recall by up to 5.9%.

So data decisions are made on transfer measurements: the extrapolated gain from a further 500k annotated images was about one point, and we did not generate them. Ranking arms by their final epoch also predicts out-of-distribution behaviour better than picking their best one (ρ = 0.77 against 0.66).

and a third arm, on how much annotation is worth adding

An audited Open Images extension of 62,589 images, mixed in at a deliberately amplified 25% relation share, was negative on every axis: about −6% on the development composite and −40.2% on its worst cell. It carries 4.33 relations per image against RA-4M's 9.03, so its unannotated pairs enter training as false negatives. Sparse annotation is worse than none, and the threshold at which that happens is set by the loss, not the data.

Deployment

20 ms per frame, and it runs in a browser

At batch 1 the cost is kernel dispatch, not parameter count. The baseline scores about 9,500 pairs among 98 boxes at 800×1333 in fp32; we score at most 128 sampled pairs among 20 boxes at 448 px in bf16.

SystemparamsA40FPS
OvSGTR Swin-T177M194.0 ms5.1
OvSGTR Swin-B237M228.5 ms4.3
RelateAnything + YOLO-World231M25.0 ms40.0

Eager PyTorch on both sides, whole system including the detector, because that is how the baseline can be timed. Compiled, ours reaches 20.3 ms (49 FPS) on an A40 and 18.1 ms on an H100; on eight CPU threads through OpenVINO, 7 FPS, agreeing with fp32 on 95.5% of top-1 predictions. Scoring 19,103 strings instead of 50 costs under a millisecond. No accuracy number on this page uses compiled inference.

browser, WASM
≈ 600 ms relation head, 80–500 ms detector
download
105 MB model, 5–27 MB per detector, then cached
calibration
two floats; ECE 0.176 → 0.004

The model exports to a single ONNX graph whose predicate bank is an input tensor, so the demo switches to any list of strings without a re-export. Weights are stored in fp16 and computed in fp32, which halves the download and runs on both ONNX Runtime Web providers.

Citation

BibTeX

@article{neau2026relateanything,
  title   = {RelateAnything: Real-Time Open-Vocabulary Relation Prediction From Any Inputs},
  author  = {Neau, Ma\"{e}lic},
  journal = {arXiv preprint},
  year    = {2026}
}