本目录提供路径驱动的 Atom0 离线测评入口。脚本本身不保存 checkpoint、数据集、norm、tokenizer、源码仓库或 Python 环境的固定路径;这些位置都在执行时显式传入,因此可用于不同机器和不同 Atom0 checkpoint。
当前按 B200 checkpoint 参数树和模型源码分为三种后端:
pi05-unified80:PI0.5、80 维统一动作空间、单 action head。pi05-ego-head:PI0.5、80 维统一动作空间、robot/ego 双 action head。fastwam-unified80:FastWAM、80 维统一动作空间、model.safetensors。
PI0.5 Legacy32 被明确排除;checkpoint 预检检测到输出维度为 32 时会直接报错。
当前“通用”特指 Atom0 Unified80 评测协议。Legacy32、PI0.7 和 StarVLA 只在此记录为不支持范围,不在本脚本中实现;它们的动作空间、归一化、horizon 和数据接口不同,不能直接复用本脚本或与 Atom0 指标混报。
双头 PI0.5 当前源码的 sample_actions() 固定使用 robot head。因此:
- robot 数据可执行 action sampling 和 flow loss;
- EgoVerse 数据只执行 flow loss;
- 对 EgoVerse 请求 action sampling 会停止并解释原因,不会误把 robot head 结果当成 ego head 结果。
本目录不复制 OpenPI、JAX、PyTorch、FastWAM、TFDS 或模型 tokenizer。
统一入口 evaluate.py 只负责 checkpoint 结构预检和启动目标环境。执行时必须提供:
--python:目标源码环境中的 Python;--repo-root:与 checkpoint 架构匹配的源码仓库根目录;--checkpoint:本次 checkpoint;- 后端自己的数据、norm、config 或 tokenizer 参数。
入口会把本目录、<repo-root>/src 和 <repo-root> 加入子进程 PYTHONPATH。这意味着脚本可放入 GitHub,而依赖仍由运行机器上的路径定位。
--python 会保留传入的 venv 符号链接,不会解析成系统 Python;否则会丢失 venv 的 site-packages。
正式测评默认使用:
- 每个 task 取 3 条 trajectory;
- 每条 trajectory 均匀取 20 个 action anchor;
- 每条 trajectory 均匀取 10 个 flow anchor;
- action sampling 比较前 8 步;
- flow 使用模型完整 horizon:PI0.5 为 config 声明值,FastWAM 为 model spec 声明值;
- seed 为 42;
- 每个 flow anchor 默认采 1 组 noise/time。
anchor 只从能够提供完整 action target 的起点中选择,不对 action target 做尾部补齐。FastWAM 的第 9 个视频帧位于 action window 末端之后;只有该视频帧越界时复用 episode 最后一帧,与既有 FastWAM 评测约定一致。
20/10 是正式测评的覆盖率与计算量折中。Piper30 的典型 trajectory 约为 370–380 帧:20 个 action anchor 的相邻间隔约为 20 帧;10 个 flow anchor 的相邻间隔约为 36 帧,而 flow horizon 为 50,已经能够连续覆盖轨迹主要阶段。增加 anchor 只提高指标对整条轨迹的覆盖率和稳定性,不会提高模型本身的预测精度。
正式测评不得通过命令行把 anchor 数覆盖为其他值。logs/run_formal_weizhongxing_19999*.sh 与其 195 Action / 117 Flow 结果是 2026-07-31 的历史 5/3 协议记录,不属于当前正式 20/10 协议;保留它们仅用于结果溯源。
--smoke 会把选择缩小为 1 task × 1 trajectory × 1 anchor。建议 smoke 时再显式指定 --mode action,避免额外编译 flow 图。
profiles/ 描述数据语义,不描述文件路径。profile 包含:
- 数据集 ID 和原生 state/action 维度;
domain(robot/ego)与action_type(joint/eef);- EEF 数据可选的固定
coordinate_frame或动态coordinate_frame_source; - 原始 TFDS 字段到三个模型相机槽的映射;
- prompt 与 prompt prefix 来源;
- action 字段;
- action 是否已经是逐帧对齐的未来 chunk。
已提供 Piper30、Piper2 profile,以及 EgoVerse actions_cartesian[100] 的模板。复制模板后只修改真实数据集的字段语义;不要把数据目录写进 profile。
Atom0 双 action head 的 Ego/Robot 路由只读取 profile 的显式 domain,不再从 dataset_id 前缀猜测。PI0.5 flow 保持模型 train 参数的静态默认值;不要通过 module_jit 显式传入 train=False,否则该 Python 布尔值会进入 JAX tracing。
Piper profile 还声明了 Direction Match 所需的:
direction_deadband:原生动作单位下的静止阈值,可为标量或与 action 同宽的数组;direction_groups:需要单独汇总的维度组。
只有确认 action chunk 表示随时间变化的原生绝对目标时才应配置这两个字段。语义未经确认的 profile 应省略它们,脚本不会为该数据集生成含义不可靠的 Direction Match。
profiles/piper30.json 以 B200 的 realworld_piper_infidata/1.1.0 为正式数据契约。该构建来自 5586 条源轨迹,split seed 为 0:
train:4927 条,来自 10 个 seen 任务;seen_test:259 条,与 train 为同一组 10 个任务,每个任务约 5%;unseen_test:400 条,其中 cola → paper box 为 300 条,green block → basket 为 100 条;- 两个 unseen 任务不得出现在 train 或 seen_test。
10 个 seen 任务的逐任务轨迹数固定在 profile 的 expected_task_episode_counts 中。除 open the drawer(43/865,4.9711%)和 pick up the green plate and put it in the dish rack(40/801,4.9938%)因整数取整略低于 5% 外,其余任务均为精确 5%。总体 train/seen 为 4927/259,即 95.0058%/4.9942%。
评测脚本不会重新切分或移动轨迹。正式运行会先要求数据目录包含 TFDS dataset_info.json 以及 B200 构建产生的 split_summary.json、split_manifest.jsonl,并校验数据集名称/版本、5586 个唯一 episode、seed、逐 split 总数、逐任务计数、unseen 集合和三个文件的内部一致性;任一项不匹配都会在模型加载前停止。运行机只需保存当前 --split 对应的完整 TFRecord shards,不要求为了 seen/unseen 离线测评同步 train shards。校验结果、各 split 的本地 shard 可用性及三个文件的 SHA-256 会写入 run_summary.json 的 split_artifacts。通过完整数据校验后,脚本才按每任务 3 条 trajectory 和 20/10 anchors 进行测评抽样。
每个 action anchor 输出:
action_mae、action_rmse和max_absolute_error;per_dimension_mae;nan_inf_rate;- profile 声明 Direction Match 语义时输出方向指标。
方向指标分成两套:
-
movement_direction_match_legacy严格复现旧 Piper 定义,比较
sign(pred[t+1]-pred[t])与sign(target[t+1]-target[t]),包括静止维度。该值只用于历史结果对比。 -
movement_direction_match_moving使用 profile 的 deadband,只统计 target 发生有效运动且 prediction/target 均有限的位置。Piper profile 还分别输出
joint_direction_match_moving和gripper_direction_match_moving。
每个 Direction Match 同时记录 matches 和 comparisons。summary.json 使用总 matches 除以总 comparisons,不会把有效运动数量差异很大的 anchor 进行无权平均。summary 中的 max_absolute_error 是所有 anchor 的全局最大值,不是每个 anchor 最大值的平均。
实际使用的 deadband 和分组会同时写入 action/summary.json 与 run_summary.json,结果离开 profile 后仍可追溯。
python evaluate.py \
--backend pi05-unified80 \
--python /path/to/environment/bin/python \
--repo-root /path/to/matching/source \
--checkpoint /path/to/checkpoint/step \
--inspect-only预检会验证:
- PI0.5 checkpoint 是完整 Orbax
params/; - FastWAM checkpoint 是 safetensors;
- 输出维度是 80;
- PI0.5 是否包含
ego_action_out_proj与所选后端一致。
python evaluate.py \
--backend pi05-unified80 \
--python /path/to/environment/bin/python \
--repo-root /path/to/matching/source \
--checkpoint /path/to/checkpoint/step \
--config-module openpi.cotrain.config \
--config-name CONFIG_NAME \
--assets-dir /path/to/config/assets \
--norm-dir /path/to/config/assets/DATASET_ID \
--dataset-dir /path/to/tfds/builder/version \
--data-profile profiles/piper30.json \
--split seen_test \
--output-dir /path/to/results--assets-dir 是传给 config.data.create() 的目录。通常 --norm-dir 必须精确等于 <assets-dir>/<profile.dataset_id>;如果 config 声明了 norm_stats_source_config,则必须等于 <assets-dir>/../<norm_stats_source_config>/<profile.dataset_id>。脚本按照 config 的真实解析规则校验该路径,还会确认 config 中存在该 dataset,并校验 norm 的 Unified80 fingerprint。
fingerprint 默认严格校验。仅允许一种已审核的历史兼容情况:旧 metadata 尚未包含、且当前值为空的 fk_eef_slots 字段;其余 mapping、delta 或 supervision 差异全部报错。
python evaluate.py \
--backend fastwam-unified80 \
--python /path/to/environment/bin/python \
--repo-root /path/to/matching/source \
--checkpoint /path/to/checkpoint/model.safetensors \
--model-spec model_specs/fastwam_unified80_b200.json \
--tokenizer-dir /path/to/local/tokenizer \
--norm-dir /path/to/norm/DATASET_ID \
--dataset-dir /path/to/tfds/builder/version \
--data-profile profiles/piper2.json \
--split seen_test \
--output-dir /path/to/resultsmodel_specs/fastwam_unified80_b200.json 只保存经过 B200 checkpoint 验证的架构参数,不包含任何权重或依赖路径。加载使用 strict=True,模型结构与 checkpoint 不一致时直接停止。
需要补充 FK EEF 槽位的 FastWAM checkpoint 使用:
--fill-fk-eef --urdf-dir /path/to/urdfPI0.5 只允许以下链路:
原生数据 → Unified80 映射 → absolute-to-delta → DispatchNormalize(一次)
模型输出 → Unnormalize(一次)→ delta-to-absolute → 原生维度
PI0.5 后端会检查 data transforms:必须恰好有一个 DispatchNormalize,且不能再有外层 Normalize。脚本没有调用 create_trained_policy(),因为该通用 policy 工厂会额外插入一次 Normalize。
FastWAM 使用显式的单次 Normalize/Unnormalize 边界。每次运行都会写出 transform_audit.json 和 run_summary.json 中的 normalization ledger;缺少或重复约定事件会报错。
action_anchor_manifest.jsonl:实际 action anchors;flow_anchor_manifest.jsonl:实际 flow anchors;action/anchor_metrics.csv:每个 anchor 的 MAE、RMSE、最大误差和 Direction Match 等;action/summary.json:action 汇总;flow_anchor_metrics.jsonl:每个 flow anchor 的损失;flow_summary.json:flow 汇总;transform_audit.json:transform、norm 和模型契约;run_summary.json:本次参数与所有汇总。
在带 NumPy 的任意 Python 环境中运行:
PYTHONPATH=. python -m unittest discover -s tests -v
python -m compileall -q .