A TensorRT FP16 / INT8 reproduction of CBCL-PR (Ayub & Wagner, IEEE TCDS 2023) on Jetson Orin Nano, plus an ablation study quantifying how feature-backbone quantization interacts with class-incremental learning, and a short research note on a hypothesized latency channel for trust collapse in human-robot interaction.
This repository is a faithful reimplementation of the original aliayub7/CBCL-PR codebase, ported to modern PyTorch and extended with TensorRT engine builds and on-device benchmarks.
Phase 1 complete: upstream audit and project scaffolding (see
docs/AUDIT.md).
Phases 2 through 7 in progress.
| Phase | Device |
|---|---|
| FP32 baseline + ablation training | NVIDIA DGX Spark (GB10 Blackwell, 128 GB unified) |
| ONNX export | DGX Spark |
| TRT FP16 / INT8 engine compile | Jetson Orin Nano Super 8GB, JetPack 6.2, TensorRT 10.3.0 |
| Latency benchmarks | Jetson Orin Nano Super 8GB |
upstream/ Untouched clone of github.com/aliayub7/CBCL-PR for reference
src/ Modern reimplementation
scripts/ Feature extraction, ONNX export, engine build, benchmark drivers
configs/ YAML configs per dataset / per quantization mode
calibration/ INT8 calibration data and scripts
benchmarks/ Latency harness output
results/ CSVs and plots checked into git
tests/ Numerical-equivalence and smoke tests
docs/ AUDIT.md, REPRODUCIBILITY.md, RESEARCH_NOTE.md
Target: a graduate student with one Jetson Orin Nano can reproduce all benchmark
numbers in under one hour. See docs/REPRODUCIBILITY.md (added in Phase 5).
All numbers measured on a Jetson Orin Nano Super 8GB at MAXN_SUPER, JetPack 6.2, TensorRT 10.3.0. Power is full-board VDD_IN sampled at 200 ms cadence during 15 s of sustained inference. Energy per image converts power and throughput into the actual deployment-relevant metric.
| Precision | Final-increment acc | p95 B=1 | p95 B=64 | Throughput B=64 | Mean board power B=64 | Energy / image B=64 |
|---|---|---|---|---|---|---|
| FP32 (Jetson TRT) | 0.6700 +/- 0.0014 | 2.71 ms | 137.9 ms | 469 img/s | 22.0 W | 46.9 mJ |
| FP16 (Jetson TRT) | 0.6704 +/- 0.0013 | 1.34 ms | 55.3 ms | 1182 img/s | 21.9 W | 18.5 mJ |
| INT8 (Jetson TRT) | 0.6719 +/- 0.0007 | 1.23 ms | 22.8 ms | 2856 img/s | 20.0 W | 7.0 mJ |
Idle baseline: ~6.9 W. FP32 vs INT8 at B=64: 6.1x throughput, 6.7x energy efficiency, identical accuracy. All three precisions also agree within statistical noise on the full per-increment forgetting trajectory across all 50 class-incremental updates, not just at the final increment. INT8 quantization on a prototype-based CL backbone is therefore a free deployment win on the Jetson Orin Nano edge platform: no measurable accuracy or forgetting-dynamics degradation, six-fold gains in throughput and energy.
At B=1 (single-image HRI, where latency dominates): INT8 = 1.23 ms vs FP32 = 2.71 ms (2.2x faster) at 13.4 W vs 20.3 W (6.9 W headroom for other workloads on a power-constrained robot).
Charts:
results/cifar100_resnet34_pareto_b64.png- latency vs accuracy at batch 64results/cifar100_resnet34_pareto_b1.png- latency vs accuracy at batch 1results/cifar100_resnet34_forgetting_curves.png- per-increment accuracy across all 50 incrementsresults/cifar100_resnet34_energy.png- energy per inference and throughput vs batch size, per precision
See docs/RESEARCH_NOTE.md for the full position paper and the proposed
Wizard-of-Oz follow-up tying this to perceived-competence latency thresholds
in human-robot interaction.
PolyForm Noncommercial 1.0.0. The upstream CBCL-PR code retains its
original MIT license, see upstream/LICENSE.
@ARTICLE{ayub_tcds_2023,
author={Ayub, Ali and Wagner, Alan R.},
journal={IEEE Transactions on Cognitive and Developmental Systems},
title={CBCL-PR: A Cognitively Inspired Model for Class-Incremental Learning in Robotics},
year={2023},
doi={10.1109/TCDS.2023.3299755}
}