TrustContact studies when a vision-language-model contact proposal should be used—or rejected—by a downstream two-person SMPL-X optimizer.
The central distinction is between semantic correctness and conditional downstream utility. A plausible human-human contact edge can improve pair geometry, have no measurable effect, or cause harmful changes in penetration, relative layout, pose, or reprojection. TrustContact therefore labels contact interventions through paired optimization and evaluates selective use/abstain policies instead of treating VLM confidence as optimization confidence.
Research status: the Contact Intervention Utility (CIU) problem and evaluation MVP are complete. A geometry risk filter showed a promising development signal, but the frozen small held-out confirmation failed its preregistered helpful-retention gate. This repository does not claim a validated final selector.
The VLM is a frozen proposal/evidence source. It does not directly predict
SMPL-X parameters and does not make the final use/abstain decision. Utility
labels come from paired O1 - O0 optimizer outcomes under an otherwise frozen
context.
Given one RGB image containing two closely interacting people:
- recover an initial pair of camera-space SMPL-X bodies with a frozen HMR initializer;
- generate a fixed bank of candidate contact edges from VLM and geometric sources;
- run paired optimization with no candidate edge (
O0) and one candidate edge (O1); - record continuous contact-gain and seven-axis safety deltas;
- label each intervention as
helpful,neutral, orharmfulusing frozen non-inferiority thresholds; - train and evaluate a lightweight risk model that can use or abstain from a proposed edge.
The utility implementation is in
src/trustcontact/utility.py, and the frozen
Gate-A thresholds are in configs/gate_a_oracle.yaml.
The 50-event Gate-A experiment found substantial contact-distance headroom, but unconditional contact failed the frozen penetration margins. A separate safety-aware Gate-A2 context passed aggregate non-inferiority on an untouched 18-event reserve cohort; the original Gate-A failure remains unchanged.
| Experiment | Main observation | Frozen decision |
|---|---|---|
| Gate-A, 50 events | O2-O0 GT-edge distance: -93.23 mm; penetration max: +13.24 mm | Fail safety |
| Gate-A2, reserve 18 | GT-edge distance: -80.86 mm; penetration max: -0.66 mm | Pass in this optimizer context |
See docs/EXPERIMENT_EXECUTION_STATUS.md.
CIU-F3 completed 156/156 K1000 optimizer cells for 32 events and 124 executable candidates. Every pair has continuous targets and a complete seven-axis safety vector.
| Candidate stratum | Count | Helpful | Neutral | Harmful |
|---|---|---|---|---|
| GT-positive | 32 | 15 | 10 | 7 |
| Nearest hard-negative | 32 | 14 | 9 | 9 |
| Distance-matched negative | 32 | 11 | 6 | 15 |
| Natural Qwen proposal | 28 | 3 | 5 | 20 |
| Total | 124 | 43 | 30 | 51 |
The natural Qwen proposals had a 71.4% harmful rate when always used. This is evidence that semantic proposals cannot be inserted into the optimizer without a downstream risk decision; it is not a claim that every VLM or optimizer has the same failure rate.
At 25% coverage on the 28 frozen Qwen proposals, the sequence-grouped DEV risk filter accepted seven proposals: two helpful, four neutral, and one harmful. Harmful rate fell from 71.4% for always-use to 14.3%, while helpful retention was 66.7%. The cluster-bootstrap interval included zero, and the accepted set was small, so this result only authorized held-out confirmation.
CSV tables and machine-readable outputs are available under
outputs/ciu_core_selector_v1.
CIU-VAL used 14 sequence-disjoint events. Qwen produced 11 executable proposals
and three explicit abstentions. Always-use had six harmful proposals out of 11.
The frozen filter accepted two proposals with zero harmful outcomes, but retained
only one of four helpful proposals (25%), below the preregistered 40% minimum.
The strict decision is therefore fail_small_heldout_confirmation.
The post-hoc ranking diagnostic is kept separate and does not relabel the frozen decision. TEST-ID and ACTION-OOD were not accessed.
configs/ Frozen experiment protocols, schemas, prompts, and gates
docs/ Execution status, asset notes, and architecture diagrams
research/ Charter, literature map, research questions, logs, and CSV ledgers
src/ Auditable TrustContact Python package
scripts/ Dataset, inference, optimization, evaluation, and audit entrypoints
tests/ Unit and integrity tests
audit/ Machine-readable gate closures and provenance records
outputs/ Curated public tables and result figures only
Datasets, model weights, environments, caches, third-party repositories, raw images, derived image overlays, and full optimizer result trees are intentionally excluded from the public repository.
The lightweight TrustContact package requires Python 3.10 or newer.
python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e .
pytest -qThe public snapshot was verified with 157 passing tests on Python 3.11.5.
Many experiment scripts additionally require licensed datasets, SMPL/SMPL-X assets, pinned third-party repositories, model checkpoints, CUDA, and the server-side directory layout recorded in the experiment protocols. The public snapshot is sufficient for source review and lightweight unit tests, but it is not a self-contained release of external assets.
- splits are grouped by event or sequence rather than random image rows;
- candidate banks and selector features are frozen before outcome access;
- forbidden post-intervention fields are audited out of selector inputs;
- manifests bind image paths, dimensions, byte counts, and SHA-256 digests;
- missing people and failed optimizer cells are preserved rather than silently removed;
- development, validation, test, construction-only, and post-hoc roles are explicitly separated;
- failed gates remain failed after diagnostic analysis.
The dynamic history is in research/research_log.md,
and the experiment ledger is in
research/tables/experiment_registry.csv.
- the positive selector claim did not pass the frozen CIU-VAL retention gate;
- main CIU results currently use Hi4D Camera04, PromptHMR GT-box interaction-off initialization, and the ProsePose HHCS optimizer context;
- licensed CHI3D raw images are not included and fresh CHI3D image baselines remain blocked;
- the current evidence does not establish cross-initializer, cross-optimizer, or cross-dataset generalization;
- Qwen fixed-edge verification and fine-grained body-part grounding failed the frozen Gate-V0 checks;
- no license has yet been assigned to the TrustContact source snapshot.
- Project charter
- Literature and field map
- Research gaps, questions, and baselines
- Roadmap and gates
- VLM contact-sensor study
- Baseline audit
- Current status and revised route
- Innovation gates
- BR1 response-probe result
- BR1 safety trajectory audit
No license is currently granted for the TrustContact source in this repository. External datasets, body models, checkpoints, and third-party code retain their original licenses and must be obtained from their official sources.


