Preprint · 2026
Any region·any relation·in real time
An image, a set of boxes, and a list of relation names as text go in. A score for every pair of boxes against every name comes out, at 20 ms per frame. The model is never given object class labels, and it runs in your browser.
Weights on Hugging Face · the corpus as RA-4M
Each graph is the 53M model's own output on a Creative Commons photo. Run it on your own images; nothing is uploaded.
One 45-second clip, three ways of drawing the regions: boxes, then instance masks, then those masks with the names hidden. The same ONNX graph as the demo throughout, handed box coordinates and pixels, never words — and a filter on calibrated log-odds holds each relation through the frames where an endpoint goes undetected.
Detection and segmentation have already moved the taxonomy out of the model and into the input. Relation prediction has not: scene-graph models still learn one corpus's 50 predicates, conditioned on object class labels. RelateAnything takes pixels, regions from any source, and a vocabulary supplied as strings. On three benchmarks that contributed no training image and a fourth evaluated zero-shot, it has 2.3–3.5× the mean recall of the strongest open-vocabulary method of comparable scale and 5–21× its rare-predicate recall, at 7.8× the speed; the six-axis composite is 40.1 against 11.8. It also leads a scene-graph model built on a 3B vision–language model on both recall metrics on all three benchmarks both were run on, at under 2% of its parameters.
Getting there meant measuring the metric first. A frequency table that never sees an image beats this model on the number leaderboards are ranked by — so the benchmark here reports six axes, two of them scored against negatives a human adjudicated.
An open-weight VLM annotates images marked with a numbered dot per box, in three passes; deterministic geometric checks then reject 11.3% of what it writes. 10,102 free-text predicates, and 9.03 relations per image against the source annotations' 5.29 on the same images and boxes.
We know of no other machine-annotated relation corpus that checks its annotator against box geometry.
Pixels and box coordinates in, one calibrated score per predicate string out. The strings are encoded by a distilled text encoder into an ℓ₂-normalised bank that no learned layer touches.
Adding a predicate means adding a row. Regions can come from any detector, or from a segmenter with no class names at all.
Transfer, precision, open vocabulary, deployment, graph quality, spatial — all scored across datasets. Each can be satisfied on its own by a different shortcut, which is why all six are reported together.
Every recall number is reported beside the shared triplet mass between training corpus and benchmark.
Interactive · nothing to download
Pick a photo and choose which relations to score. The model scored 263 predicate strings on each of these photos once; toggling a word re-runs the demo's decode on those stored scores. Switch the region source to FastSAM and the objects lose their names entirely — the graph is computed the same way.
Results · cross-dataset, no object labels
Six axes, each scored on benchmarks that contributed no training image, each against the strongest system that can be run on it. Pick an axis; the chart carries the numbers and the notes carry the caveats.
Transfer. Closed-vocabulary recall, ground-truth boxes, four sources of differing shared triplet mass.Alone: gameable by corpus match.
Precision. Federated AP against adjudicated negatives, the only axis that sees a confident false positive.Alone: invariant to uniform score depression.
Open vocabulary. Full training vocabulary deployed, synonym-tolerant matching: does the model mean the right relation?Alone: rewards head collapse onto on / in.
Deployment. Detection-mode graphs on a shared detector, against the measured pair-recall ceiling.Alone: bounded by the detector rather than the model.
Graph quality. A VLM judge, one relation at a time, no ground truth in the prompt; each acceptance credited with its surprisal.Alone: shown whole graphs, the same judge prefers short repetitive ones. That verdict is reported too.
Spatial. SpatialSense: adversarial, balanced true/false, where knowing that on is common does not help.Alone: a boxes-only baseline scores 68.8 without the image.
The composite is the chance-corrected harmonic mean over A1, A2, A4, A5 and A6, with A1 entering as F1@50 averaged over its four benchmarks. It spans only the axes a baseline can be run on, so OvSGTR alone carries it; A3 is reported and left out, because the one baseline runnable there is scored through a matcher whose choice moves the result by more than the two systems differ.
OvR-SGG holds 15 of VG150's 50 predicates out of training and scores in detection mode. We reproduce all twelve published numbers to within 0.35 points, then retrained with those 15 strings and their synonym groups removed.
| Method | backbone / boxes | B+N R@50 | Novel R@50 | B+N mR@50 | Novel mR@50 | classes hit |
|---|---|---|---|---|---|---|
| VS3 | Swin-T | 15.6 | 0.0 | n/r | n/r | n/r |
| OvSGTR | Swin-T | 20.5 | 13.5 | n/r | n/r | n/r |
| RAHP | Swin-T | 20.5 | 15.6 | n/r | n/r | n/r |
| OvSGTR + MegaSG | Swin-T | 25.4 | 17.0 | n/r | n/r | n/r |
| OvSGTR | Swin-B | 22.9 | 16.4 | n/r | n/r | n/r |
| INOVA | Swin-B | 24.8 | 20.0 | n/r | n/r | n/r |
| OvSGTR re-run here | Swin-T | 20.4 | 13.2 | 4.1 | 1.9 | 13/50 3/15 novel |
| RelateAnything novel 15 held out | their boxes | 22.4 | 11.8 | 14.9 | 7.2 | 40/50 8/15 novel |
| RelateAnything zero-shot tower, novel set seen | their boxes | 27.3 | 22.8 | n/r | n/r | n/r |
| RelateAnything released, novel set seen | their boxes | 27.9 | 24.1 | n/r | n/r | n/r |
On the leaderboard's micro metric the held-out model is mid-table, and we claim no lead there. On mean recall over the same predictions it is nearly four times the baseline, answering with 40 of the 50 predicates against 13. The split explains both: the 15 held-out predicates carry 55% of the test relations, and on, of and in account for 93% of that. The last two rows have those 15 strings in training, so their Novel columns do not satisfy the protocol; the gap to them, −5.5 and −12.3 R@50, measures how much of novel performance here is supervision on those three words. Mean recall and the class count are measured here for the two systems we run and reported by no published row (n/r).
Exact-string matching penalises an open-vocabulary model for its own synonyms: to the right of scores zero against right of. With synonym-tolerant matching over the full training vocabulary, the correct relation is the model's median first choice on two of the three benchmarks.
| Source | R@50 | mR@50 | rare | MRR | median rank |
|---|---|---|---|---|---|
| VG150 | 56.0 | 34.5 | 39.6 | 0.68 | 1 / 19,103 |
| PSG | 30.5 | 28.3 | 20.8 | 0.39 | 10 / 19,103 |
| IndoorVG | 53.3 | 34.6 | 34.6 | 0.65 | 1 / 19,103 |
For scale, the published closed-vocabulary ceilings under PredCls are 68.2 R@50 (PE-Net) and 37.0 mR@50 (VCTree+IETrans+Rwt) on VG150, and 41.7 mR@50 on PSG (DSFlash-L, under the corrected protocol of Lorenz et al.) — from specialists trained on that benchmark, with ground-truth object labels and a fixed head.
A system that answers in free text can be scored on A3 by construction, which is the one comparison OvSGTR cannot enter: its vocabulary arrives as a single caption, about 150 strings. That admits ROBIN-3B, a scene-graph model built on a 3B vision–language model, and general multimodal models prompted for the same output. Below is one set of generations per model read several ways, so a system's rows differ only in how its free text is mapped onto the benchmark's words.
| Model | matcher | R@20 | R@50 | mR@20 | mR@50 | pairs reached |
|---|---|---|---|---|---|---|
| PSG test, 2,179 images, ground-truth regions for every model, one scorer | ||||||
| RelateAnything 53M | exact string | 12.3 | 13.6 | 12.4 | 13.0 | 99.7 |
| RelateAnything | synonym, τ = 0.60 | 31.7 | 36.8 | 29.1 | 31.3 | 99.7 |
| ROBIN-3B | exact string | 25.8 | 26.2 | 19.9 | 20.0 | 45.6 |
| ROBIN-3B | SBERT argmax its own | 30.4 | 35.7 | 21.1 | 22.9 | 77.4 |
| ROBIN-3B | synonym, τ = 0.60 | 35.1 | 37.9 | 24.1 | 25.1 | 65.6 |
| Qwen3-VL-32B prompted | synonym, τ = 0.60 | 15.0 | 15.6 | 5.1 | 5.2 | 35.5 |
| Qwen3-VL-8B prompted | synonym, τ = 0.60 | 10.2 | 10.2 | 3.3 | 3.3 | 23.6 |
| InternVL3.5-8B prompted | synonym, τ = 0.60 | 9.5 | 9.6 | 4.5 | 4.5 | 23.1 |
| GLM-4.6V-Flash prompted | synonym, τ = 0.60 | 8.9 | 9.1 | 2.1 | 2.2 | 23.4 |
The matcher is a scoring convention whose effect exceeds the difference it is used to measure: mapping free text onto the benchmark's words is worth +11.7 R@50 to ROBIN and +23.2 to us, since a model answering from 19,103 strings is the one exact matching penalises hardest. It also decides the mean-recall ordering — ROBIN leads on exact strings, 20.0 against 13.0, and we lead under every synonym-tolerant matcher, 31.3 against 25.1. Two things survive it. One is pairs reached, the share of annotated pairs a system names at all: 99.7% for us, 45.6–77.4% for ROBIN, 23.1–35.5% for the prompted models, and a pair never named cannot carry the right predicate. The other is what happens off PSG: run on three benchmarks, each under its own protocol, we lead ROBIN on both metrics on all three, by 1.9× on VG150 micro recall and 2.2× on IndoorVG against 1.1× on PSG.
The model
A visual path and a text path that meet at a cosine. The figure is the paper's own; the chips under it take the blocks one at a time.
The standard protocol supplies ground-truth boxes and their classes, so a model can lean on the prior that a ⟨person, bicycle⟩ pair is probably riding. Ours sees pixels and box coordinates only, which is the situation when boxes come from a live detector. Measured: 87–93% of the semantic logit's variance comes from pair context and 0.1% from object identity, and scoring against another image's features costs 44–68% of micro accuracy.
At this vocabulary size the supervision is positive-unlabeled: a ⟨man, horse⟩ pair annotated riding is also sitting on, and scoring that as a negative teaches the model to suppress correct answers. Each negative is therefore discounted by the estimated probability that it also holds — fitted on the 706k training pairs carrying more than one annotation — while annotated predicates and directional inverses keep full weight, since separating above from below is the one signal that must not be softened.
A distilled 512-d encoder maps each string into the head's space, and nothing there is trained after the bank is built, so a string never seen in training is treated exactly like one that was. It has to be distilled because contrastive encoders are direction-blind: the teacher puts above and below at cosine 0.95 and the two left/right forms at 0.99, indistinguishable from its synonym pairs at 0.96. No visual model regressing onto such targets can separate them, whoever trains it.
The corpus
Human relation annotation is expensive and annotators agree with each other poorly. RA-4M generates annotations at scale and then filters them against box geometry.
A gate fires only where box geometry logically constrains the predicate, so a rejection is a true negative up to box error, and what geometry cannot constrain passes through unchecked and is counted rather than guessed at. Contact predicates must touch; containment and proximity are checked against the boxes; direction is repaired rather than rejected, the roles swapped to match the geometry with the annotator's own wording kept. Because the annotator prefers the left and upper object as subject, each verified directional relation is then restated from the other endpoint with probability one half — roles swapped, predicate replaced by its inverse — which leaves the corpus direction-balanced, a property the model depends on and no source corpus has.
| Corpus | images | relations | rel/img | predicates | entropy |
|---|---|---|---|---|---|
| RA-4M | 474,413 | 4,282,531 | 9.03 | 10,102 | 3.99 |
| MegaSG source annotations, same images | 474,420 | 2,510,905 | 5.29 | 94 | 2.56 |
| raw Visual Genome | 108,077 | 2,315,906 | 22.10 | 36,549 | 3.81 |
| after the leakage filter (secondary source) | 40,615 | 758,547 | 18.70 | 17,742 | — |
Against the annotations it replaces, on the same images and the same boxes: 1.7× denser, 107× the vocabulary, 1.4 nats more predicate entropy, and 73 of the 94 source predicate classes reproduced as exact strings. Mean object degree rises from 1.88 to 3.20 and the share of images whose graph is a single connected component from 70.9% to 82.6%. The largest unverifiable class is gaze: about 22% of looking at and watching annotations have disjoint boxes.

Measurement
Four results that changed how we read every number above.
The freq baseline looks up the most frequent training predicate for a ⟨subject, object⟩ category pair. It reads the ground-truth object labels and never sees the image. It wins on per-edge accuracy on all three benchmarks — 68.4 against 57.7 on VG150 — and loses on the per-predicate average of the same predictions, 18.9 against 35.1. A metric a lookup table can win does not measure relation understanding, and it is the metric that orders leaderboards.
Top-1 over ground-truth pairs and boxes, the model being the tower trained without any HICO-DET share. 6.9–12.8% of edges are ones we get right and it gets wrong, so pixels are clearly doing something.
on between a person and a horse and on between a book and a table are different acts; two corpora can agree on the string and never agree on the pair it is asserted of. Shared triplet mass is the share of a training corpus's relation instances whose ⟨subject category, predicate, object category⟩ triple the benchmark also annotates. It collapses when the object categories are charged for, and it collapses unevenly: on VG150 the baseline's fine-tuning corpus scores 90.9% by that measure, against our 12.8%. The confound is concentrated in-domain, and it is very large there.
| Share of the corpus's relations matched on | VG150 | PSG | IndoorVG | Haystack |
|---|---|---|---|---|
| our released training mixture · 19,103 predicate strings | ||||
| the predicate string | 53.4 | 38.7 | 49.5 | 38.7 |
| both object categories | 44.4 | 36.7 | 26.8 | 30.7 |
| the whole triple | 12.8 | 10.7 | 6.6 | 4.9 |
| VG150 train · 50 predicates · the baseline's fine-tuning corpus | ||||
| the predicate string | 100.0 | 57.3 | 95.7 | 57.3 |
| both object categories | 100.0 | 22.1 | 12.3 | 7.1 |
| the whole triple | 90.9 | 8.6 | 10.4 | 0.6 |
No image and no annotation is shared with any benchmark; leakage is excluded separately, by image identity. The statistic is predictive — one arm of ours with a larger share of raw Visual Genome reached 54.3 R@50 on VG150, the best zero-shot figure we know of, and was also the worst model we trained on every tail metric.
The standard detection-mode matcher has no assignment constraint, so a ground-truth object covered by d detections gives d chances at the same relation. On VG150 test the mean is 2.99. Adding the constraint costs the published baseline 2.45 R@50 points; upgrading its backbone from Swin-T to Swin-B is worth 2.39.
Detection-mode recall is bounded by pair recall, which is quadratic in object recall. With the detector weights fixed, moving only the confidence threshold and the box cap shifts that ceiling by 19 points on VG150. The spread between all published methods on that leaderboard is 4.3 points.
One cheap fix covers most of this: state the matcher, the detector's threshold and box cap, and the pair-recall ceiling they imply, next to every detection-mode number.
Every ablation was read twice, once on the corpus's own validation split and once on benchmarks the model never trained on. Doubling the training corpus is worth +15% micro and +56% macro recall in-domain, and +2.5–3.8% and +9.9–11.9% under transfer: the extra images teach the corpus's annotation style, which the in-domain split shares by construction. One arm changes sign outright — removing an automatically derived left/right excess improves in-domain metrics at every epoch and lowers transfer macro recall by up to 5.9%.
So data decisions are made on transfer measurements: the extrapolated gain from a further 500k annotated images was about one point, and we did not generate them. Ranking arms by their final epoch also predicts out-of-distribution behaviour better than picking their best one (ρ = 0.77 against 0.66).
An audited Open Images extension of 62,589 images, mixed in at a deliberately amplified 25% relation share, was negative on every axis: about −6% on the development composite and −40.2% on its worst cell. It carries 4.33 relations per image against RA-4M's 9.03, so its unannotated pairs enter training as false negatives. Sparse annotation is worse than none, and the threshold at which that happens is set by the loss, not the data.
Deployment
At batch 1 the cost is kernel dispatch, not parameter count. The baseline scores about 9,500 pairs among 98 boxes at 800×1333 in fp32; we score at most 128 sampled pairs among 20 boxes at 448 px in bf16.
| System | params | A40 | FPS |
|---|---|---|---|
| OvSGTR Swin-T | 177M | 194.0 ms | 5.1 |
| OvSGTR Swin-B | 237M | 228.5 ms | 4.3 |
| RelateAnything + YOLO-World | 231M | 25.0 ms | 40.0 |
Eager PyTorch on both sides, whole system including the detector, because that is how the baseline can be timed. Compiled, ours reaches 20.3 ms (49 FPS) on an A40 and 18.1 ms on an H100; on eight CPU threads through OpenVINO, 7 FPS, agreeing with fp32 on 95.5% of top-1 predictions. Scoring 19,103 strings instead of 50 costs under a millisecond. No accuracy number on this page uses compiled inference.
The model exports to a single ONNX graph whose predicate bank is an input tensor, so the demo switches to any list of strings without a re-export. Weights are stored in fp16 and computed in fp32, which halves the download and runs on both ONNX Runtime Web providers.
Citation
@article{neau2026relateanything,
title = {RelateAnything: Real-Time Open-Vocabulary Relation Prediction From Any Inputs},
author = {Neau, Ma\"{e}lic},
journal = {arXiv preprint},
year = {2026}
}