Skip to content

Repository files navigation

TrustContact

TrustContact studies when a vision-language-model contact proposal should be used—or rejected—by a downstream two-person SMPL-X optimizer.

The central distinction is between semantic correctness and conditional downstream utility. A plausible human-human contact edge can improve pair geometry, have no measurable effect, or cause harmful changes in penetration, relative layout, pose, or reprojection. TrustContact therefore labels contact interventions through paired optimization and evaluates selective use/abstain policies instead of treating VLM confidence as optimization confidence.

Research status: the Contact Intervention Utility (CIU) problem and evaluation MVP are complete. A geometry risk filter showed a promising development signal, but the frozen small held-out confirmation failed its preregistered helpful-retention gate. This repository does not claim a validated final selector.

中文项目说明见 中文项目报告、 项目总览 和 研究工作区。

Architecture

TrustContact architecture

The VLM is a frozen proposal/evidence source. It does not directly predict SMPL-X parameters and does not make the final use/abstain decision. Utility labels come from paired O1 - O0 optimizer outcomes under an otherwise frozen context.

Core task

Given one RGB image containing two closely interacting people:

  1. recover an initial pair of camera-space SMPL-X bodies with a frozen HMR initializer;
  2. generate a fixed bank of candidate contact edges from VLM and geometric sources;
  3. run paired optimization with no candidate edge (O0) and one candidate edge (O1);
  4. record continuous contact-gain and seven-axis safety deltas;
  5. label each intervention as helpful, neutral, or harmful using frozen non-inferiority thresholds;
  6. train and evaluate a lightweight risk model that can use or abstain from a proposed edge.

The utility implementation is in src/trustcontact/utility.py, and the frozen Gate-A thresholds are in configs/gate_a_oracle.yaml.

Main evidence

Contact headroom and safety

The 50-event Gate-A experiment found substantial contact-distance headroom, but unconditional contact failed the frozen penetration margins. A separate safety-aware Gate-A2 context passed aggregate non-inferiority on an untouched 18-event reserve cohort; the original Gate-A failure remains unchanged.

Experiment Main observation Frozen decision
Gate-A, 50 events O2-O0 GT-edge distance: -93.23 mm; penetration max: +13.24 mm Fail safety
Gate-A2, reserve 18 GT-edge distance: -80.86 mm; penetration max: -0.66 mm Pass in this optimizer context

See docs/EXPERIMENT_EXECUTION_STATUS.md.

CIU-DEV full outcomes

CIU-F3 completed 156/156 K1000 optimizer cells for 32 events and 124 executable candidates. Every pair has continuous targets and a complete seven-axis safety vector.

Candidate stratum Count Helpful Neutral Harmful
GT-positive 32 15 10 7
Nearest hard-negative 32 14 9 9
Distance-matched negative 32 11 6 15
Natural Qwen proposal 28 3 5 20
Total 124 43 30 51

The natural Qwen proposals had a 71.4% harmful rate when always used. This is evidence that semantic proposals cannot be inserted into the optimizer without a downstream risk decision; it is not a claim that every VLM or optimizer has the same failure rate.

Development selector signal

At 25% coverage on the 28 frozen Qwen proposals, the sequence-grouped DEV risk filter accepted seven proposals: two helpful, four neutral, and one harmful. Harmful rate fell from 71.4% for always-use to 14.3%, while helpful retention was 66.7%. The cluster-bootstrap interval included zero, and the accepted set was small, so this result only authorized held-out confirmation.

CIU-DEV Qwen proposal filter curve

CIU-DEV coverage and safety curve

CSV tables and machine-readable outputs are available under outputs/ciu_core_selector_v1.

Frozen small held-out confirmation

CIU-VAL used 14 sequence-disjoint events. Qwen produced 11 executable proposals and three explicit abstentions. Always-use had six harmful proposals out of 11. The frozen filter accepted two proposals with zero harmful outcomes, but retained only one of four helpful proposals (25%), below the preregistered 40% minimum. The strict decision is therefore fail_small_heldout_confirmation.

Frozen CIU-VAL confirmation

The post-hoc ranking diagnostic is kept separate and does not relabel the frozen decision. TEST-ID and ACTION-OOD were not accessed.

Repository layout

configs/       Frozen experiment protocols, schemas, prompts, and gates
docs/          Execution status, asset notes, and architecture diagrams
research/      Charter, literature map, research questions, logs, and CSV ledgers
src/           Auditable TrustContact Python package
scripts/       Dataset, inference, optimization, evaluation, and audit entrypoints
tests/         Unit and integrity tests
audit/         Machine-readable gate closures and provenance records
outputs/       Curated public tables and result figures only

Datasets, model weights, environments, caches, third-party repositories, raw images, derived image overlays, and full optimizer result trees are intentionally excluded from the public repository.

Installation

The lightweight TrustContact package requires Python 3.10 or newer.

python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e .
pytest -q

The public snapshot was verified with 157 passing tests on Python 3.11.5.

Many experiment scripts additionally require licensed datasets, SMPL/SMPL-X assets, pinned third-party repositories, model checkpoints, CUDA, and the server-side directory layout recorded in the experiment protocols. The public snapshot is sufficient for source review and lightweight unit tests, but it is not a self-contained release of external assets.

Reproducibility and integrity

  • splits are grouped by event or sequence rather than random image rows;
  • candidate banks and selector features are frozen before outcome access;
  • forbidden post-intervention fields are audited out of selector inputs;
  • manifests bind image paths, dimensions, byte counts, and SHA-256 digests;
  • missing people and failed optimizer cells are preserved rather than silently removed;
  • development, validation, test, construction-only, and post-hoc roles are explicitly separated;
  • failed gates remain failed after diagnostic analysis.

The dynamic history is in research/research_log.md, and the experiment ledger is in research/tables/experiment_registry.csv.

Current limitations

  • the positive selector claim did not pass the frozen CIU-VAL retention gate;
  • main CIU results currently use Hi4D Camera04, PromptHMR GT-box interaction-off initialization, and the ProsePose HHCS optimizer context;
  • licensed CHI3D raw images are not included and fresh CHI3D image baselines remain blocked;
  • the current evidence does not establish cross-initializer, cross-optimizer, or cross-dataset generalization;
  • Qwen fixed-edge verification and fine-grained body-part grounding failed the frozen Gate-V0 checks;
  • no license has yet been assigned to the TrustContact source snapshot.

Documentation

License and external assets

No license is currently granted for the TrustContact source in this repository. External datasets, body models, checkpoints, and third-party code retain their original licenses and must be obtained from their official sources.

About

视觉语言模型约束接触

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages