Date: 2026-08-26
Status source: local workspace plus live server audit
GATE-A2 is complete and passes in the audited safer-optimizer context:
hhcmap=0.10 plus hhcgen=0.03. The preregistered original
GATE-A-MAIN-50 decision remains fail, because its unconditional contact
intervention did not satisfy the frozen penetration margins. Gate-A2 is a new
control experiment; it does not retroactively relabel Gate-A.
The result decomposes into two scientifically useful findings:
- contact headroom: supported under the cached BEV initializer and
hhcmapweight 0.1; - unconditional contact safety: not supported in the same context.
- contact headroom with aggregate non-inferiority safety: supported under
the audited
hhcgen=0.03context on the untouched 18-event reserve cohort.
No VLM-value or selector-performance claim has been tested yet.
The Hi4D portion of the baseline gate is now complete under a frozen 100-frame, 25-sequence, single-camera manifest. PromptHMR detector-native off/on and the ECCV 2024 Multi-HMR 896L checkpoint have both run on byte-identical images and been evaluated with a shared SMPL-24 joint adapter. This authorizes a small, locked Gate-V0 sensor audit; it does not authorize large-scale VLM inference or selector training.
- Server root:
/root/autodl-tmp/vlmcontact - GPU: NVIDIA GeForce RTX 5090, 32607 MiB
- Environment: Python 3.10, PyTorch 2.7.1+cu128
- PyTorch3D: 0.7.9, CUDA operators verified
- SMPL-X neutral asset:
/root/autodl-tmp/vlmcontact/assets/body_models/smplx/SMPLX_NEUTRAL.npz - SMPL-X SHA256:
376021446ddc86e99acacd795182bbef903e61d33b76b9d8b359c2b0865bd992 - Local handoff archive:
models_smplx_v1_1.zip; archive contains neutral, male and female SMPL-X assets - Raw CHI3D image count on the server: 0
- Hi4D licensed archive: 95,560 images, 191,120 masks, train/test sequence-disjoint; exact audit complete
- Frozen Hi4D baseline manifest:
HI4D-BASELINE-TEST100-V1, 100 rows from 25 test sequences, Camera04, four evenly spaced frames per sequence - Current Gate-A2 background compute jobs: none after completion
- Multi-HMR:
651fb411e1cbcc626aaa5f38805ecab9cc891f7a - BUDDI:
041cdbc0866f84d8fb21952448fec8564f9b4679 - ProsePose:
df635f7d50117e8b35ea44387f6371cec2f942a4 - PromptHMR:
3b566b7dbb28ce506c7ea972c18693f4c705ce8c
BEV, BUDDI, ProsePose and PromptHMR have not yet been accepted as fully reproduced paper baselines; an official-image engineering smoke or use of cached data/loss code is not equivalent to a unified CHI3D reproduction.
- Local TrustContact tests: 27/27 passed with the project source on
PYTHONPATHand unrelated pytest plugin auto-loading disabled. - Remote Gate-A2 runners, evaluator and analyzer compiled and executed end to end; local TrustContact tests remain 10/10 passed.
- Multi-HMR official ECCV checkpoint
multiHMR_896_L.ptis pinned by SHA-256c1b55933...02ca; 100/100 Hi4D inferences completed. DINOv2 source is loaded offline from a hash-recorded archive, avoiding an unfrozen torch-hub fetch. - PromptHMR official checkpoint and config were pinned by SHA-256 and loaded with the unmodified official strict loader. A controlled official-image interaction-off/on smoke passed; this is not a CHI3D accuracy result.
- BUDDI modules and PyTorch3D CUDA path import and execute.
- ProsePose processed CHI3D val/test files are present.
- Gate-A scripts, evaluator, event manifest and bootstrap analysis execute end to end.
- A second fail-closed pre-inference gate now verifies the frozen manifest lock plus every image's path boundary, byte count, SHA-256 and dimensions. PromptHMR and Multi-HMR are not authorized to start if this audit is incomplete or any frozen input has changed.
- Resumable manifest runners are staged for PromptHMR box-only interaction off/on and Multi-HMR. Every event artifact is bound to the frozen image hash and full run-configuration hash; mismatched artifacts are never overwritten.
scripts/run_chi3d_unified_baselines.shnow provides an explicitpreflightmode and an opt-inrunmode. Its server preflight stopped at 0/50 exactly as required, before loading either baseline model.- Local workspace is not currently a Git repository; commit-level provenance for TrustContact itself is therefore still missing.
- 118 eligible CHI3D validation events
- Metric and identity plumbing only;
scientific_result=false
- One CHI3D event
- Confirmed that released
hhcdistminignores the supplied 75x75 map - Froze the contact-map-aware
hhcmapimplementation for Gate-A
- 10 events x 2 conditions x 3 repeats = 60/60 runs
- No failures; empirical repeat variation was effectively zero
- Authorized the disjoint main cohort only
- 50 disjoint action-stratified events
- O0 no contact, O1 one GT edge, O2 all GT edges
- 150/150 result cells; zero failures
- Mean runtime: 56.74 seconds per optimizer run
O2 minus O0:
| Metric | Mean delta | Event-bootstrap 95% CI | Decision |
|---|---|---|---|
| GT-edge distance | -93.23 mm | [-113.30, -74.50] | contact headroom passes |
| Relative translation error | -58.10 mm | [-77.70, -39.45] | passes |
| MPJPE | -0.81 mm | [-1.36, -0.30] | passes |
| Translation-aligned PVE | -1.17 mm | [-2.41, -0.13] | passes |
| Contact F1 at 30 mm | +0.025 | [0.008, 0.047] | diagnostic only |
| Penetration maximum | +13.24 mm | [7.41, 19.53] | fails 10 mm margin |
| Penetration mean | +3.55 mm | [1.53, 5.62] | fails 3 mm margin |
| Reprojection error | +0.046 px | [0.015, 0.079] | passes |
The 20-event development cohort compared O1 weights 0.03, 0.10 and 0.30 against O0: 80/80 cells, zero failures. Both 0.03 and 0.10 passed development; the frozen gain rule selected 0.10. On the independent 20-event confirmation cohort, 40/40 cells completed and contact gain remained strong, but penetration safety failed:
| Metric, O1(0.10)-O0 | Mean delta | 95% CI | Decision |
|---|---|---|---|
| GT-edge distance | -41.58 mm | [-70.75, -20.09] | primary passes |
| Relative translation error | -17.90 mm | [-34.30, -6.45] | passes |
| Penetration maximum | +5.97 mm | [1.45, 12.79] | fails 10 mm upper bound |
| Penetration mean | +1.91 mm | [-0.03, 4.23] | fails 3 mm upper bound |
confirmation_pass=false; this cohort was not rerun or reused for selection.
The released ProsePose GeneralContactLoss (hhcgen) was audited before use.
On development record 1053 at weight 0.10, the weighted loss was 0.468 and the
translation-gradient norm was 45.81. A paired four-step smoke reduced O1
penetration max from 64.50 to 58.12 mm and mean from 18.71 to 16.60 mm.
With contact weight fixed at 0.10, the 20-event development sweep completed
120/120 cells for hhcgen weights 0.03/0.10/0.30. All three were eligible;
the frozen within-5-mm tie rule selected 0.03. That configuration was then run
once on all 18 untouched reserve events: 36/36 cells, zero failures.
| Metric, O1-O0 on reserve | Mean delta | Event-bootstrap 95% CI | Decision |
|---|---|---|---|
| GT-edge distance | -80.86 mm | [-118.67, -47.46] | primary passes; 18/18 improve |
| Relative translation error | -44.44 mm | [-75.12, -21.45] | passes |
| MPJPE | -0.02 mm | [-0.90, 1.16] | passes 3 mm margin |
| PA-MPJPE | +1.08 mm | [0.13, 2.25] | passes 3 mm margin |
| Translation-aligned PVE | -0.63 mm | [-1.40, 0.15] | passes 3 mm margin |
| Penetration maximum | -0.66 mm | [-10.32, 6.21] | passes 10 mm margin |
| Penetration mean | +0.67 mm | [-1.59, 2.60] | passes 3 mm margin |
| Reprojection error | +0.021 px | [0.009, 0.035] | passes 0.25 px margin |
Final Gate-A2 decision: confirmation_pass=true for this optimizer context.
Eight helpful, two neutral and eight harmful event labels remain on reserve;
aggregate safety passing therefore does not remove the selective-utility
problem, but per-event predictability is still untested.
- Official repository commit:
3b566b7dbb28ce506c7ea972c18693f4c705ce8c; - official checkpoint SHA-256:
a3ef04ef8a12c3068682b03c62c95f8959cd8554424e105c63bf97f6c8e97e99; - official config SHA-256:
31142aaac12a7ef60adba77c85addffad1ea466a6555717b24c3fbed079a2447; - isolated RTX-5090 compatibility environment: Python 3.10, PyTorch 2.7.1+cu128 and xFormers 0.0.31;
- the released xFormers kernels failed on
sm_120, so the parameter-free attention operator was replaced at runtime by PyTorch SDPA. The official tracked source stayed clean and checkpoint loading remainedstrict=True; - identical official image, identical top-two YOLO boxes and no mask prompt; only cross-person attention changed from off to on;
- warm forward times: 53.26 ms off and 56.63 ms on;
- paired mean output displacement: vertices 41.13 mm, body joints 42.21 mm, translations 43.01 mm.
This only proves that the released switch is executable and non-degenerate.
It does not show that interaction-on is more accurate. The released evaluator
also has a CHI3D naming mismatch (CHI3D_TEST in the entry script versus no
CHI3D constructor branch), so the unified CHI3D comparison requires an audited
adapter rather than an unsupported claim of official-paper reproduction.
CHI3D-RAW-IMAGE-FREEZE-V1 is implemented and tested fail-closed. It targets
the already frozen 50-event primary-view cohort and records relative path,
dimensions, bytes and SHA-256. On both local and server roots it resolved 0/50
because licensed raw images are absent; it wrote an audit report and did not
write chi3d_baseline_main50_v1.csv.
- Exact archive/annotation audit passed: 75 train and 25 test sequences are disjoint; all 95,560 annotated frames and 191,120 person masks are present; every frame has two identities, eight synchronized cameras and SMPL-6890 one-based contact correspondence values (zero is the no-contact sentinel).
HI4D-BASELINE-TEST100-V1freezes 25 test sequences x four frames from Camera04. Every extracted image is bound by path, dimensions, bytes and SHA-256. Statistical intervals resample sequences, not individual frames.- PromptHMR detector-native: 97/100 two-person coverage; the three preserved
failures are
cheers32/000096,dance27/000061, anddance28/000049. - PromptHMR GT-box mechanism control: 100/100 coverage.
- Multi-HMR 896L: 100/100 inference completion, 93/100 two-person coverage; seven one-person outputs remain explicit failures. Median forward time is 42.84 ms. The 2026 Anny checkpoint remains an efficiency/engineering comparison only, not the scientific ECCV baseline.
Shared detector-native SMPL-24 results (mean; sequence-cluster 95% CI):
| Method | Coverage | Root MPJPE | PA-MPJPE | Pair PA-MPJPE | Relative pelvis |
|---|---|---|---|---|---|
| PromptHMR interaction off | 97% | 37.53 [20.61, 64.27] | 20.34 [14.62, 27.36] | 50.28 [38.76, 63.45] | 147.05 [103.18, 206.06] |
| PromptHMR interaction on | 97% | 36.98 [20.05, 63.85] | 20.05 [14.37, 27.05] | 49.96 [38.81, 62.74] | 146.30 [103.46, 204.24] |
| Multi-HMR 896L | 93% | 70.11 [59.32, 84.48] | 54.86 [45.67, 65.60] | 85.89 [71.62, 103.93] | 159.43 [120.02, 205.05] |
Geometry means are conditional on successful two-person output; coverage is reported separately and missing people are never silently dropped. These are TrustContact's shared SMPL-24 results, not copied paper numbers and not an exact H36M-14/PVE reproduction.
The GT-box interaction control isolates PromptHMR's cross-person switch. On minus off gives root MPJPE -0.378 mm, CI [-0.709, -0.028], but pair PA-MPJPE +1.140 mm, CI [+0.115, +2.208], and relative-pelvis error +3.692 mm, CI [+0.626, +6.796]. Thus interaction-on improves individual pose slightly while harming pair layout in this controlled cohort. This is direct evidence for context-dependent intervention utility, not evidence that a TrustContact selector can predict it.
The 2026-08-25 exporter applies the thresholds already frozen before the main pilot; it does not tune thresholds after seeing the results.
| Intervention | Helpful | Neutral | Harmful |
|---|---|---|---|
| O1: one GT edge | 21 | 6 | 23 |
| O2: all GT edges | 15 | 5 | 30 |
Most harmful labels are caused by penetration. This establishes outcome heterogeneity and motivates selective use, but it does not prove that pre-optimization VLM or geometry features can predict those labels.
- Exact PromptHMR paper-protocol H36M-14/PVE evaluation remains blocked by the unavailable Google Drive evaluation annotations and licensed gendered SMPL assets. The transparent SMPL-24 shared evaluator is complete.
- Fresh CHI3D image baselines remain blocked because licensed CHI3D raw images are absent; Hi4D now provides the runnable cross-dataset baseline track.
- Gate V0, fixed-candidate semantic residual, selector training and external generalization are all untested.
- Gate-A2 is limited to cached BEV initialization and CHI3D validation subject s02; it is not a cross-subject or cross-initializer claim.
- The local workspace is not under version control.
- Freeze the 75-region ontology/version and define the Hi4D 20-image prompt debugging subset plus a disjoint locked Gate-V0 audit subset by sequence.
- Run only the Qwen3-VL-8B-Instruct formatting/identity/contact-sensor smoke on the 20 debug images; cache every JSON with image/model/prompt/seed hashes.
- Audit JSON validity, A/B identity, left-right errors and candidate precision before authorizing the locked model comparison.
- Continue pursuing licensed CHI3D raw images and exact PromptHMR evaluation assets in parallel; do not change any frozen Gate-A cohort.
Large-scale VLM inference and selector training remain unauthorized.