Skip to content

Latest commit

 

History

History
283 lines (229 loc) · 13.9 KB

File metadata and controls

283 lines (229 loc) · 13.9 KB

TrustContact execution status

Date: 2026-08-26
Status source: local workspace plus live server audit

Current gate

GATE-A2 is complete and passes in the audited safer-optimizer context: hhcmap=0.10 plus hhcgen=0.03. The preregistered original GATE-A-MAIN-50 decision remains fail, because its unconditional contact intervention did not satisfy the frozen penetration margins. Gate-A2 is a new control experiment; it does not retroactively relabel Gate-A.

The result decomposes into two scientifically useful findings:

  • contact headroom: supported under the cached BEV initializer and hhcmap weight 0.1;
  • unconditional contact safety: not supported in the same context.
  • contact headroom with aggregate non-inferiority safety: supported under the audited hhcgen=0.03 context on the untouched 18-event reserve cohort.

No VLM-value or selector-performance claim has been tested yet.

The Hi4D portion of the baseline gate is now complete under a frozen 100-frame, 25-sequence, single-camera manifest. PromptHMR detector-native off/on and the ECCV 2024 Multi-HMR 896L checkpoint have both run on byte-identical images and been evaluated with a shared SMPL-24 joint adapter. This authorizes a small, locked Gate-V0 sensor audit; it does not authorize large-scale VLM inference or selector training.

Confirmed resources

  • Server root: /root/autodl-tmp/vlmcontact
  • GPU: NVIDIA GeForce RTX 5090, 32607 MiB
  • Environment: Python 3.10, PyTorch 2.7.1+cu128
  • PyTorch3D: 0.7.9, CUDA operators verified
  • SMPL-X neutral asset: /root/autodl-tmp/vlmcontact/assets/body_models/smplx/SMPLX_NEUTRAL.npz
  • SMPL-X SHA256: 376021446ddc86e99acacd795182bbef903e61d33b76b9d8b359c2b0865bd992
  • Local handoff archive: models_smplx_v1_1.zip; archive contains neutral, male and female SMPL-X assets
  • Raw CHI3D image count on the server: 0
  • Hi4D licensed archive: 95,560 images, 191,120 masks, train/test sequence-disjoint; exact audit complete
  • Frozen Hi4D baseline manifest: HI4D-BASELINE-TEST100-V1, 100 rows from 25 test sequences, Camera04, four evenly spaced frames per sequence
  • Current Gate-A2 background compute jobs: none after completion

Pinned repositories

  • Multi-HMR: 651fb411e1cbcc626aaa5f38805ecab9cc891f7a
  • BUDDI: 041cdbc0866f84d8fb21952448fec8564f9b4679
  • ProsePose: df635f7d50117e8b35ea44387f6371cec2f942a4
  • PromptHMR: 3b566b7dbb28ce506c7ea972c18693f4c705ce8c

BEV, BUDDI, ProsePose and PromptHMR have not yet been accepted as fully reproduced paper baselines; an official-image engineering smoke or use of cached data/loss code is not equivalent to a unified CHI3D reproduction.

Verified engineering state

  • Local TrustContact tests: 27/27 passed with the project source on PYTHONPATH and unrelated pytest plugin auto-loading disabled.
  • Remote Gate-A2 runners, evaluator and analyzer compiled and executed end to end; local TrustContact tests remain 10/10 passed.
  • Multi-HMR official ECCV checkpoint multiHMR_896_L.pt is pinned by SHA-256 c1b55933...02ca; 100/100 Hi4D inferences completed. DINOv2 source is loaded offline from a hash-recorded archive, avoiding an unfrozen torch-hub fetch.
  • PromptHMR official checkpoint and config were pinned by SHA-256 and loaded with the unmodified official strict loader. A controlled official-image interaction-off/on smoke passed; this is not a CHI3D accuracy result.
  • BUDDI modules and PyTorch3D CUDA path import and execute.
  • ProsePose processed CHI3D val/test files are present.
  • Gate-A scripts, evaluator, event manifest and bootstrap analysis execute end to end.
  • A second fail-closed pre-inference gate now verifies the frozen manifest lock plus every image's path boundary, byte count, SHA-256 and dimensions. PromptHMR and Multi-HMR are not authorized to start if this audit is incomplete or any frozen input has changed.
  • Resumable manifest runners are staged for PromptHMR box-only interaction off/on and Multi-HMR. Every event artifact is bound to the frozen image hash and full run-configuration hash; mismatched artifacts are never overwritten.
  • scripts/run_chi3d_unified_baselines.sh now provides an explicit preflight mode and an opt-in run mode. Its server preflight stopped at 0/50 exactly as required, before loading either baseline model.
  • Local workspace is not currently a Git repository; commit-level provenance for TrustContact itself is therefore still missing.

Completed experiments

Cached-initializer preflight

  • 118 eligible CHI3D validation events
  • Metric and identity plumbing only; scientific_result=false

Contact-aware convergence smoke

  • One CHI3D event
  • Confirmed that released hhcdistmin ignores the supplied 75x75 map
  • Froze the contact-map-aware hhcmap implementation for Gate-A

Optimizer-noise pilot

  • 10 events x 2 conditions x 3 repeats = 60/60 runs
  • No failures; empirical repeat variation was effectively zero
  • Authorized the disjoint main cohort only

Main Gate-A

  • 50 disjoint action-stratified events
  • O0 no contact, O1 one GT edge, O2 all GT edges
  • 150/150 result cells; zero failures
  • Mean runtime: 56.74 seconds per optimizer run

O2 minus O0:

Metric Mean delta Event-bootstrap 95% CI Decision
GT-edge distance -93.23 mm [-113.30, -74.50] contact headroom passes
Relative translation error -58.10 mm [-77.70, -39.45] passes
MPJPE -0.81 mm [-1.36, -0.30] passes
Translation-aligned PVE -1.17 mm [-2.41, -0.13] passes
Contact F1 at 30 mm +0.025 [0.008, 0.047] diagnostic only
Penetration maximum +13.24 mm [7.41, 19.53] fails 10 mm margin
Penetration mean +3.55 mm [1.53, 5.62] fails 3 mm margin
Reprojection error +0.046 px [0.015, 0.079] passes

Gate-A2 weight-only control

The 20-event development cohort compared O1 weights 0.03, 0.10 and 0.30 against O0: 80/80 cells, zero failures. Both 0.03 and 0.10 passed development; the frozen gain rule selected 0.10. On the independent 20-event confirmation cohort, 40/40 cells completed and contact gain remained strong, but penetration safety failed:

Metric, O1(0.10)-O0 Mean delta 95% CI Decision
GT-edge distance -41.58 mm [-70.75, -20.09] primary passes
Relative translation error -17.90 mm [-34.30, -6.45] passes
Penetration maximum +5.97 mm [1.45, 12.79] fails 10 mm upper bound
Penetration mean +1.91 mm [-0.03, 4.23] fails 3 mm upper bound

confirmation_pass=false; this cohort was not rerun or reused for selection.

Gate-A2 audited safety-loss control

The released ProsePose GeneralContactLoss (hhcgen) was audited before use. On development record 1053 at weight 0.10, the weighted loss was 0.468 and the translation-gradient norm was 45.81. A paired four-step smoke reduced O1 penetration max from 64.50 to 58.12 mm and mean from 18.71 to 16.60 mm.

With contact weight fixed at 0.10, the 20-event development sweep completed 120/120 cells for hhcgen weights 0.03/0.10/0.30. All three were eligible; the frozen within-5-mm tie rule selected 0.03. That configuration was then run once on all 18 untouched reserve events: 36/36 cells, zero failures.

Metric, O1-O0 on reserve Mean delta Event-bootstrap 95% CI Decision
GT-edge distance -80.86 mm [-118.67, -47.46] primary passes; 18/18 improve
Relative translation error -44.44 mm [-75.12, -21.45] passes
MPJPE -0.02 mm [-0.90, 1.16] passes 3 mm margin
PA-MPJPE +1.08 mm [0.13, 2.25] passes 3 mm margin
Translation-aligned PVE -0.63 mm [-1.40, 0.15] passes 3 mm margin
Penetration maximum -0.66 mm [-10.32, 6.21] passes 10 mm margin
Penetration mean +0.67 mm [-1.59, 2.60] passes 3 mm margin
Reprojection error +0.021 px [0.009, 0.035] passes 0.25 px margin

Final Gate-A2 decision: confirmation_pass=true for this optimizer context. Eight helpful, two neutral and eight harmful event labels remain on reserve; aggregate safety passing therefore does not remove the selective-utility problem, but per-event predictability is still untested.

PromptHMR interaction-off/on engineering smoke

  • Official repository commit: 3b566b7dbb28ce506c7ea972c18693f4c705ce8c;
  • official checkpoint SHA-256: a3ef04ef8a12c3068682b03c62c95f8959cd8554424e105c63bf97f6c8e97e99;
  • official config SHA-256: 31142aaac12a7ef60adba77c85addffad1ea466a6555717b24c3fbed079a2447;
  • isolated RTX-5090 compatibility environment: Python 3.10, PyTorch 2.7.1+cu128 and xFormers 0.0.31;
  • the released xFormers kernels failed on sm_120, so the parameter-free attention operator was replaced at runtime by PyTorch SDPA. The official tracked source stayed clean and checkpoint loading remained strict=True;
  • identical official image, identical top-two YOLO boxes and no mask prompt; only cross-person attention changed from off to on;
  • warm forward times: 53.26 ms off and 56.63 ms on;
  • paired mean output displacement: vertices 41.13 mm, body joints 42.21 mm, translations 43.01 mm.

This only proves that the released switch is executable and non-degenerate. It does not show that interaction-on is more accurate. The released evaluator also has a CHI3D naming mismatch (CHI3D_TEST in the entry script versus no CHI3D constructor branch), so the unified CHI3D comparison requires an audited adapter rather than an unsupported claim of official-paper reproduction.

Raw-image manifest preflight

CHI3D-RAW-IMAGE-FREEZE-V1 is implemented and tested fail-closed. It targets the already frozen 50-event primary-view cohort and records relative path, dimensions, bytes and SHA-256. On both local and server roots it resolved 0/50 because licensed raw images are absent; it wrote an audit report and did not write chi3d_baseline_main50_v1.csv.

Hi4D archive, manifest and full baseline gate

  • Exact archive/annotation audit passed: 75 train and 25 test sequences are disjoint; all 95,560 annotated frames and 191,120 person masks are present; every frame has two identities, eight synchronized cameras and SMPL-6890 one-based contact correspondence values (zero is the no-contact sentinel).
  • HI4D-BASELINE-TEST100-V1 freezes 25 test sequences x four frames from Camera04. Every extracted image is bound by path, dimensions, bytes and SHA-256. Statistical intervals resample sequences, not individual frames.
  • PromptHMR detector-native: 97/100 two-person coverage; the three preserved failures are cheers32/000096, dance27/000061, and dance28/000049.
  • PromptHMR GT-box mechanism control: 100/100 coverage.
  • Multi-HMR 896L: 100/100 inference completion, 93/100 two-person coverage; seven one-person outputs remain explicit failures. Median forward time is 42.84 ms. The 2026 Anny checkpoint remains an efficiency/engineering comparison only, not the scientific ECCV baseline.

Shared detector-native SMPL-24 results (mean; sequence-cluster 95% CI):

Method Coverage Root MPJPE PA-MPJPE Pair PA-MPJPE Relative pelvis
PromptHMR interaction off 97% 37.53 [20.61, 64.27] 20.34 [14.62, 27.36] 50.28 [38.76, 63.45] 147.05 [103.18, 206.06]
PromptHMR interaction on 97% 36.98 [20.05, 63.85] 20.05 [14.37, 27.05] 49.96 [38.81, 62.74] 146.30 [103.46, 204.24]
Multi-HMR 896L 93% 70.11 [59.32, 84.48] 54.86 [45.67, 65.60] 85.89 [71.62, 103.93] 159.43 [120.02, 205.05]

Geometry means are conditional on successful two-person output; coverage is reported separately and missing people are never silently dropped. These are TrustContact's shared SMPL-24 results, not copied paper numbers and not an exact H36M-14/PVE reproduction.

The GT-box interaction control isolates PromptHMR's cross-person switch. On minus off gives root MPJPE -0.378 mm, CI [-0.709, -0.028], but pair PA-MPJPE +1.140 mm, CI [+0.115, +2.208], and relative-pelvis error +3.692 mm, CI [+0.626, +6.796]. Thus interaction-on improves individual pose slightly while harming pair layout in this controlled cohort. This is direct evidence for context-dependent intervention utility, not evidence that a TrustContact selector can predict it.

Frozen utility-label diagnostic

The 2026-08-25 exporter applies the thresholds already frozen before the main pilot; it does not tune thresholds after seeing the results.

Intervention Helpful Neutral Harmful
O1: one GT edge 21 6 23
O2: all GT edges 15 5 30

Most harmful labels are caused by penetration. This establishes outcome heterogeneity and motivates selective use, but it does not prove that pre-optimization VLM or geometry features can predict those labels.

Active blockers

  1. Exact PromptHMR paper-protocol H36M-14/PVE evaluation remains blocked by the unavailable Google Drive evaluation annotations and licensed gendered SMPL assets. The transparent SMPL-24 shared evaluator is complete.
  2. Fresh CHI3D image baselines remain blocked because licensed CHI3D raw images are absent; Hi4D now provides the runnable cross-dataset baseline track.
  3. Gate V0, fixed-candidate semantic residual, selector training and external generalization are all untested.
  4. Gate-A2 is limited to cached BEV initialization and CHI3D validation subject s02; it is not a cross-subject or cross-initializer claim.
  5. The local workspace is not under version control.

Next authorized stage

  1. Freeze the 75-region ontology/version and define the Hi4D 20-image prompt debugging subset plus a disjoint locked Gate-V0 audit subset by sequence.
  2. Run only the Qwen3-VL-8B-Instruct formatting/identity/contact-sensor smoke on the 20 debug images; cache every JSON with image/model/prompt/seed hashes.
  3. Audit JSON validity, A/B identity, left-right errors and candidate precision before authorizing the locked model comparison.
  4. Continue pursuing licensed CHI3D raw images and exact PromptHMR evaluation assets in parallel; do not change any frozen Gate-A cohort.

Large-scale VLM inference and selector training remain unauthorized.