This repository studies a deliberately adversarial SVG-editing setting: raw markup contains IDs and duplicate element tags, but removes geometric facts. A language-only model is asked to identify the element for an instruction such as “the smallest object” or “the top-left object.” It cannot reliably infer those properties from the supplied markup. We measure that information gap, then add a coordinate-free, render-derived Relative Visual Signature (RVS).
cd svg-grounding
python -m venv .venv && source .venv/bin/activate
make install
make reproduce
make testThis writes 24 deterministic scenes under data/benchmark/ and evaluation artifacts under results/latest/. Set baseline.mode: local_llm in configs/eval.yaml to use Qwen/Qwen2.5-0.5B-Instruct (weights download on first use). The default blind baseline is intentional: it is a deterministic no-geometry control that makes the information limitation reproducible offline.
Each scene has 6–8 primitives, duplicate tags, a common fill, and an opaque scene.svg given to the baseline. render.svg and render.png are evaluation assets generated from withheld geometry. Annotations provide the single ground-truth target for spatial and size/ordinal instruction families.
RVS does not expose coordinates, dimensions, or areas. It converts rendered geometry to a table of qualitative region, horizontal/vertical ordinal rank, descending size rank, nearest neighbour, and shape class. The included deterministic interpreter is an auditable offline reader; format_rvs also provides a compact prompt representation for an SLM replacement.
configs/ reproducible YAML settings
src/svg_grounding/ generator, baseline, RVS, evaluation
data/benchmark/ generated SVG/PNG/JSON assets (after generation)
results/latest/ CSV, JSON, Markdown, RVS tables, figures
tests/ invariant and grounding tests
report.md concise research report
results/latest/predictions.csv has every case and outcome. metrics.json includes overall and family accuracy, precision/recall/F1, compact confusion matrices, per-scene statistics, and failure categories. Figures pair ground-truth and baseline-prediction highlights.
Synthetic relations are intentionally clean and the RVS interpreter understands the supplied instruction grammar. This is evidence that a qualitative representation can solve this controlled identification task; it does not establish general SVG editing competence, visual understanding, or performance of a particular pretrained model. See report.md for negative findings and interpretation guidance.