# RelateAnything — local deployment & webcam demo Run the relation model on your own machine: an open-vocabulary detector proposes boxes, the relation head predicts open-vocabulary relations between them, and you get a live scene graph — one merged graph, or **two graphs** (spatial + semantic) from the same forward pass. ``` webcam ─▶ YOLO-World (RA-4M 497 objects) ─▶ boxes ─▶ RelSGG (predicates) ─▶ scene graph(s) ▲ rebuilt locally (AGPL) ▲ per-model dist bundle ``` Both vocabularies are swappable — that is the whole point. See [`deploy/vocab.py`](vocab.py) for the curated selections; every **number** (thresholds, recall, alpha) lives in per-model generated artifacts, never in code. **ONNX is the shipping path**: torch-free on the target, numpy + onnxruntime only. The torch path still exists for development. **Browser demo**: the same graphs also run fully client-side (ONNX Runtime Web, WebGPU/WASM) at https://maelic.github.io/RelateAnything_demo/ — repo `Maelic/RelateAnything_demo`, model files regenerated by `deploy/web/export_web_models.py`. --- ## 1. Build a model's distribution (once, on the training box) One driver runs the whole chain — ONNX export with parity check → per-checkpoint threshold calibration → predicate bank: ```bash python deploy/build_release.py --only relsgg-vits16plus # -> deploy/dist/relsgg-vits16plus/{relateanything.onnx, relateanything.json, # predicate_bank.npz, thresholds.json} ``` Models are declared in [`release_manifest.json`](release_manifest.json). The individual steps remain runnable by hand (`export_onnx.py`, `calibrate_thresholds.py`, `build_predicate_bank.py`) — the driver only sequences them. Three invariants the tooling enforces rather than trusts: * **Text space** — the bank encoder is resolved from the checkpoint's own args and cross-checked against the training embeddings (min cosine ≥ 0.999 on shared strings) before a bank is written. * **Thresholds are per-checkpoint** — score scales do not transfer between models, so uncalibrated bank rows are NaN, loudly, never borrowed numbers. * **Provenance** — `relateanything.json` records run name, git SHA, backbone, text-student sha256, merge state, and the measured ONNX parity delta. ## 2. Detector (rebuilt locally, never redistributed) The detectors derive from ultralytics (**AGPL-3.0**) and are excluded from every release artifact (see `THIRD_PARTY_NOTICES.md`). Rebuild locally: ```bash python deploy/reparam_detector.py # re-parameterize to RA-4M-497 python deploy/export_detector_onnx.py --size s --imgsz 640 mkdir -p deploy/dist/detector-local cp checkpoints/detectors/yolov8s-worldv2_megasg497.onnx deploy/dist/detector-local/detector.onnx cp checkpoints/detectors/yolov8s-worldv2_megasg497.json deploy/dist/detector-local/detector.json ``` ## 3. Run the demo ```bash pip install -r deploy/dist/requirements.txt # numpy, opencv, onnxruntime python deploy/demo_webcam.py --dist deploy/dist/relsgg-vits16plus # webcam python deploy/demo_webcam.py --dist deploy/dist/relsgg-vits16plus --image photo.jpg python deploy/demo_webcam.py --dist deploy/dist/relsgg-vits16plus --image photo.jpg --decompose ``` Live keys: `q` quit · `[` `]` threshold · `p` vocabulary preset · `+`/`-` triplet count · `b` box-confidence weighting · **`g` two-graph mode** (spatial edges orange, semantic white). ## 4. Dynamic vocabulary & thresholds The ONNX graph takes `W [V,768]` and `alpha [V]` as *inputs*: swapping the predicate vocabulary is slicing different rows out of the model's `predicate_bank.npz` — no re-export, no text encoder on the target. Adding a string the bank lacks requires re-running `build_predicate_bank.py` where the text student lives. Per-predicate operating points come from the bank's calibrated `thr` array (measured on that checkpoint's own scores by `calibrate_thresholds.py`); `ThresholdConfig` values you set explicitly always win. All decode behavior (global floor, per-predicate overrides, pair weighting, two-graph split) is host-side numpy in [`postprocess.py`](postprocess.py) — mutate between frames freely. ## 5. Two-graph mode `--decompose` renders a **spatial** and a **semantic** graph from the same forward pass, using the bank's `is_spatial` vector (type-stratified graph constraint — each pair may contribute one edge per stream, so a pair can legitimately hold `on` and `holding` at once). Banks built before this schema do not offer the feature, and say so rather than degrading silently. ## 6. Torch backend (development) ```bash python deploy/demo_webcam.py --backend torch --checkpoint /model.pth ``` ## TensorRT (NVIDIA GPU) The native TensorRT 10 backend builds engines from the same released ONNX graphs. It shares preprocessing, vocabulary selection, calibration and decoding with the ONNX backend, including `--decompose`. PyTorch provides CUDA buffers and streams; inference runs in TensorRT. Use a CUDA-enabled PyTorch install and an NVIDIA driver compatible with both PyTorch and TensorRT. From the repository root on a Linux x86-64 NVIDIA machine: ```bash pip install -e ".[tensorrt,hub]" # Download the graph into the bundle already containing the bank and sidecar. hf download maelic/relsgg-vits16plus relateanything.onnx \ --local-dir deploy/dist/relsgg-vits16plus python deploy/export_tensorrt.py \ --onnx deploy/dist/relsgg-vits16plus/relateanything.onnx --check # After rebuilding your detector locally (section 2): python deploy/export_tensorrt.py \ --onnx deploy/dist/detector-local/detector.onnx --check python deploy/demo_webcam.py --dist deploy/dist/relsgg-vits16plus \ --backend tensorrt --device cuda --image photo.jpg --decompose ``` Each build writes `_trt.engine` and `_trt.json`, retaining the source metadata and calibration and recording the GPU, TensorRT version, ONNX hash and shape profiles. Engines are specific to the GPU/platform and TensorRT version; build them on the deployment machine, and rebuild after changing the graph or TensorRT. The source ONNX and its sidecar remain unchanged. On Jetson, install TensorRT 10 and CUDA-enabled PyTorch through NVIDIA's platform packages instead of the `tensorrt` pip extra; Jetson has not been validated here. The initial backend uses **FP32 with TF32 disabled**. Reduced precision can change pair selection and needs separate accuracy validation. The batch size, image size and padded box count are fixed to the bundle's contract (normally 1 / 448 / 32). The actual region count may be zero through `max_boxes`. The predicate count remains dynamic from 1 through the bank size; set `--max-vocab N` when building for a larger bank. Changing predicates within that profile does not rebuild the engine. Masks use the existing exported boxes-only path; segmentation can still provide those boxes. Use just the relation engine with your own region source: ```python from deploy.trt_runtime import TensorRTRelationHead from deploy.postprocess import ThresholdConfig, decode head = TensorRTRelationHead( "deploy/dist/relsgg-vits16plus/relateanything_trt.engine", "deploy/dist/relsgg-vits16plus/predicate_bank.npz", device="cuda", ) head.set_predicates(["on", "holding", "beside"]) pred, pair, sub, obj, valid = head(frame_bgr, boxes_xyxy) cfg = ThresholdConfig(threshold=0.5, calib_a=head.contract.calib_a, calib_b=head.contract.calib_b) triplets = decode(pred, pair, sub, obj, valid, head.predicates, cfg, boxes_xyxy=boxes_xyxy) ``` `--check` compares valid pair sets and raw logits with ONNX Runtime across empty/singleton inputs, several box counts, and small/default/full vocabulary selections. Invalid padding is ignored and outputs are aligned by subject/object pair because TopK ties can be reordered. Add `--check-images image1.jpg image2.jpg` to include real images (with generated boxes); these checks verify numerical agreement, not benchmark accuracy. A failed check exits with an error and does not record a successful parity report. Use `--bench` in the demo to measure latency on your own system; the published A40 figures are PyTorch measurements, not TensorRT results. All three released checkpoints have been checked on an RTX 3080 Laptop GPU with TensorRT 10.16.1.11. The latest comparison uses local exports of all three released checkpoints, including the sparse-pair ONNX CUDA padding correction. See the [family benchmark](../docs/benchmarks/README.md) for numerical checks, median and p95 latency, and commands to export and measure each model. The standalone ViT-S+ relation API also matched merged and decomposed decoded graphs after swapping between 3, 243 and 1 predicates. Jetson remains unvalidated. The [end-to-end benchmark](../docs/benchmarks/end-to-end.md) adds COCO YOLO26, YOLO-World or YOLOE detection and records detector, relation, decoding and total latency. It reports the same three checkpoints and four relation backends, with raw samples, box counts, numerical checks and thermal telemetry. --- ## 7. The project reel (`make_reel.py`) `assets/reel/hero.gif` is a rendered artifact, not a screen recording. (The README's hero is now a video reel over moving footage — see section 8.) It makes a single claim visible — the model reads *regions*, and nothing else — by changing only where the regions come from, three times: | shot | detector | drawn | |---|---|---| | `boxes` | YOLO-World v2-S · MEGASG-497 | boxes + class names | | `masks` | YOLOE-11s prompt-free · 4,585 classes | instance masks + class names | | `unnamed` | FastSAM-s · class-agnostic | instance masks, **no names at all** | ```bash pip install -e ".[deploy,reel]" python deploy/make_reel.py --models ../relate-anything/demo/models --out assets/reel # -> assets/reel/{reel.mp4, hero.gif, poster.jpg} ``` `--models` is the browser demo's export directory: the segmentation detectors are AGPL-3.0 ultralytics derivatives and are not redistributed here (see [`THIRD_PARTY_NOTICES.md`](../THIRD_PARTY_NOTICES.md)). The box detector comes from `deploy/dist/detector-local/` when it is present. The shot list, and a note per shot on why it draws the number of edges it draws, is [`assets/reel/shots.json`](../assets/reel/shots.json). Two things worth knowing before changing it: * **What runs is the exported graph**, which takes `image` and `boxes` only — the same one the browser runs. `RelateAnything.predict(..., masks=)` in the torch API rasterises masks into the head; that is a different path and the reel does not use it. In the mask shots the segmenter chooses the regions and the masks are what you see, while the boxes around them are what is scored. * **Which edges get drawn is a rule, not a preference** (`select_edges`): the model's ranking in order, minus one edge per pair, one per predicate, and anything whose endpoints are too small to see. `deploy/seg_detector.py` additionally collapses regions that are the same object under two names, which a 4,585-class vocabulary produces constantly (`persian cat` / `feline`). --- ## 8. The video reel (`render_video.py`) The README's hero is thirty seconds over five Creative Commons clips, six seconds each. It exists to show the thing a still cannot: that the graph stays put while the scene moves. ```bash pip install -e ".[deploy,reel,detector]" # detector is AGPL: the mask shots need it python deploy/fetch_footage.py --from_shots assets/reel_video/shots.json --out footage python deploy/render_video.py --reel assets/reel_video/shots.json --out hero.mp4 \ --width 1280 --fps 30 --stride 2 --crf 18 --max_edges 6 --max_spatial 2 ``` The footage is not redistributed here; `fetch_footage.py` downloads it from Commons and writes a `credits.json` beside it. `--from_shots` re-fetches exactly the clips the shot list names, by their Commons page URL, so the reel is reproducible. Two filters do the work that a per-frame model cannot: * **A Kalman filter per track** on `(cx, cy, w, h)`, with ByteTrack-style two-stage IoU association. Boxes stop jittering without lagging behind a moving subject, which one EMA cannot do — it has a single knob for two jobs. * **A Kalman filter per relation**, on the calibrated log-odds from `ScoreContract.fuse`. The distinction it draws is between a relation that was *measured low* and one that was *not measurable* because an endpoint went undetected: the first is evidence against, the second is no evidence at all. Treating them alike is what makes a per-frame graph blink. On a 180-frame clip this cut blinks by 39% and raised the mean run of an edge from 21 to 32 frames. Pass `--no_belief` to turn it off and see the difference. Why the output is 1280x720: four of the five sources are native 720p, so a larger canvas would upscale them rather than reveal anything. Only the guitarist clip is 4K, and it downscales. Crops are 16:9 by construction — a crop of any other shape gets pillarboxed into the frame, which was how an early cut ended up with black bars down both sides.