diff --git a/README.md b/README.md index 6b6fcfb84..80a9ea9bf 100644 --- a/README.md +++ b/README.md @@ -27,6 +27,15 @@ RPent framework +## Benchmark Results + +Compare RPent with reference methods on LIBERO, LIBERO-PRO, RoboCasa365 Target50, and RoboTwin C2R. Rankings apply to the methods and evaluation coverage shown; see [Benchmark Results](https://rpent.readthedocs.io/en/latest/rst_source/benchmarks.html) for suite results and model configurations. + +Codex / GPT-6 Astra / low / reasoning: **92.63% Overall (741/800)** across all eight LIBERO-PRO suites. See the [suite results and memory-batch explanation](https://rpent.readthedocs.io/en/latest/rst_source/benchmarks.html#libero-pro-astra-memory), including the separately frozen Long and Spatial/Object/Goal memory batches. + +[![RPent success-rate comparisons across four benchmarks](https://cdn.jsdelivr.net/gh/RLinf/misc@c3b9b5d4ffa360a8324c5b7aa510e1ed0876aa43/rpent/benchmarks/leaderboard-en-light.png)](https://rpent.readthedocs.io/en/latest/rst_source/benchmarks.html) + + ## Who Should Consider Using RPent? RPent is built for four kinds of users: @@ -38,6 +47,7 @@ RPent is built for four kinds of users: ## What's NEW! +- [2026/09] 🔥 Added an interactive leaderboard and consolidated benchmark results for LIBERO, LIBERO-PRO, RoboCasa365, and RoboTwin, with model comparisons and suite-level results. Explore [Benchmark Results](https://rpent.readthedocs.io/en/latest/rst_source/benchmarks.html). - [2026/09] 🔥 RPent supports Franka single-arm and dual-arm real-robot extensions. Doc: [Franka](https://rpent.readthedocs.io/en/latest/rst_source/usage/franka.html) · [Dual Franka](https://rpent.readthedocs.io/en/latest/rst_source/usage/dual_franka.html). - [2026/08] 🔥 RPent supports RoboCasa with RLDX-1 as manipulation model. See the [RoboCasa setup and Target50 guide](https://rpent.readthedocs.io/en/latest/rst_source/usage/robocasa.html). - [2026/08] 🔥 RPent supports the non-reasoning mode, which reduces average execution time by ~40%. diff --git a/README.zh-CN.md b/README.zh-CN.md index e45334ce8..800b6f524 100644 --- a/README.zh-CN.md +++ b/README.zh-CN.md @@ -27,6 +27,15 @@ RPent framework +## 基准测试结果 + +对比 RPent 与参考方法在 LIBERO、LIBERO-PRO、RoboCasa365 Target50 和 RoboTwin C2R 上的成功率。排名仅限图中方法及评测范围;套件成绩和模型配置见[基准测试结果](https://rpent.readthedocs.io/zh-cn/latest/rst_source/benchmarks.html)。 + +Codex / GPT-6 Astra / low / reasoning 已完成全部八套 LIBERO-PRO,**Overall 92.63%(741/800)**。详见 [套件汇总与 memory 批次说明](https://rpent.readthedocs.io/zh-cn/latest/rst_source/benchmarks.html#libero-pro-astra-memory),其中 Long 与 Spatial/Object/Goal 分别使用各自冻结的 memory 批次。 + +[![RPent 四项基准成功率对比](https://cdn.jsdelivr.net/gh/RLinf/misc@c3b9b5d4ffa360a8324c5b7aa510e1ed0876aa43/rpent/benchmarks/leaderboard-zh-light.png)](https://rpent.readthedocs.io/zh-cn/latest/rst_source/benchmarks.html) + + ## 适用用户 RPent 面向以下四类用户: @@ -38,6 +47,7 @@ RPent 面向以下四类用户: ## 最新动态 +- [2026/09] 🔥 新增交互式排行榜和基准测试结果汇总,覆盖 LIBERO、LIBERO-PRO、RoboCasa365 与 RoboTwin,提供模型对比及套件汇总。查看[基准测试结果](https://rpent.readthedocs.io/zh-cn/latest/rst_source/benchmarks.html)。 - [2026/08] 🔥 支持 RoboCasa,使用 RLDX-1 作为操作模型。参见 [RoboCasa 安装与 Target50 指南](https://rpent.readthedocs.io/zh-cn/latest/rst_source/usage/robocasa.html)。 - [2026/08] 🔥 新增非推理(non-reasoning)模式,平均执行时间降低约 40%。 - [2026/08] 🔥 支持 LIBERO 探索模式。文档:[LIBERO 探索模式](https://rpent.readthedocs.io/zh-cn/latest/rst_source/usage/libero.html#memory)。 diff --git a/docs/source-en/index.rst b/docs/source-en/index.rst index b40a05a80..3460db8a3 100644 --- a/docs/source-en/index.rst +++ b/docs/source-en/index.rst @@ -74,6 +74,7 @@ Welcome to RPent Overview Installation Quick Start + Benchmark Results .. toctree:: :maxdepth: 2 diff --git a/docs/source-en/rst_source/benchmarks.rst b/docs/source-en/rst_source/benchmarks.rst new file mode 100644 index 000000000..c9a203b7f --- /dev/null +++ b/docs/source-en/rst_source/benchmarks.rst @@ -0,0 +1,24 @@ +:html_theme.sidebar_secondary.remove: + +.. _benchmark-results: +.. _benchmark-leaderboard: +.. _leaderboard: + +RPent Leaderboard +================= + +.. raw:: html + + + + + +
+
+

LIBERO-PRO

Method / modelSuccess rate
Codex / GPT-6 Astra / low / reasoning92.63%
Claude Code / Opus-4.7 / max.reasoning82.4%
RPent Flash Mode72.63%
Codex / GPT-5.5 / xhigh / reasoning72.1%
ASPIRE61.36%
π_RLinf50.0%
π0.511.0%
AtomVLA6.3%
X-VLA3.8%
MolmoAct1.5%
π00.3%

GPT-6 Astra: 92.63% (741/800). RPent Flash Mode / Molmo2-8B: 72.63% (581/800).

All methods & reported scores

Method / modelOverallSpatial TaskSpatial SwapObject TaskObject SwapGoal TaskGoal SwapLong TaskLong Swap
GPT-6 Astra92.63%100%98%100%99%88%99%85%72%
Opus-4.7 / max.reasoning82.4%
RPent Flash Mode / Molmo2-8B72.63%79.00%70.00%86.00%93.00%74.00%65.00%60.00%54.00%
GPT-5.572.1%81.0%69.0%94.0%91.0%75.0%66.0%52.0%49.0%
ASPIRE61.36%60.0%51.0%95.0%98.0%45.0%81.0%38.3%22.6%
π_RLinf50.0%42.0%59.0%71.0%78.0%45.0%42.0%49.0%14.0%
π0.511.0%1.0%20.0%1.0%17.0%2.0%38.0%1.0%8.0%
AtomVLA6.3%1.0%16.0%0.0%10.0%11.0%2.0%9.0%1.0%
X-VLA3.8%0.0%0.0%8.0%2.0%9.0%1.0%10.0%0.0%
MolmoAct1.5%0.0%0.0%0.0%6.0%0.0%0.0%6.0%0.0%
π00.3%0.0%0.0%0.0%2.0%0.0%0.0%0.0%0.0%
Cap-X14.0%12.0%18.0%22.0%17.0%26.0%
RATS31.0%29.0%63.0%61.0%36.0%43.0%

ASPIRE: Long Task and Long Swap use zero-shot transfer from the LIBERO-90 skill library.

GPT-6 Astra: Long Task/Swap and the other six suites use separate memory-file snapshots frozen after their respective exploration phases, with no updates during evaluation. Overall combines two non-overlapping batches: Long 157/200 plus the other suites 584/600, giving 741/800 (92.63%); the 800 episodes do not share a single memory snapshot.

+

LIBERO

Method / modelSuccess rate
AtomVLA97.0%
Claude Code / Opus-4.7 / max.reasoning96.0%
π_RLinf95.3%
π094.2%
NORA79.5%
OpenVLA76.5%
+

RoboCasa365 · Target50

Method / modelSuccess rate
Codex / GPT-6 Astra / low / reasoning59.20%
Xiaomi-Robotics-157.4%
Codex / GPT-5.5 / xhigh / reasoning57.1%
Claude Code / Opus-4.7 / max.reasoning48.6%
WorldDreamer35.3%
RLDX-130.0%
π0.516.9%
π014.8%

RoboCasa365 · Target50 · All methods & reported scores

Method / modelOverallAtomic-SeenComposite-SeenComposite-Unseen
GPT-6 Astra59.20%87.78%43.75%42.50%
Xiaomi-Robotics-157.4%80.2%57.1%32.1%
GPT-5.557.1%92.0%61.0%13.8%
Opus-4.7 / max.reasoning48.6%79.4%47.5%15.0%
WorldDreamer35.3%66.3%26.7%9.0%
RLDX-130.0%60.0%21.3%5.0%
π0.516.9%39.6%7.1%1.2%
π014.8%34.6%6.1%1.1%
+

RoboTwin

Method / modelSuccess rate
Codex / GPT-5.5 / xhigh / reasoning62.4%
Claude Code / Opus-4.7 / max.reasoning58.4%
LingBot-VLA50.4%
π0.547.9%
GR00T-N1.720.7%
StarVLA10.6%
+
+
diff --git a/docs/source-en/rst_source/overview.rst b/docs/source-en/rst_source/overview.rst index d00a84cff..1a8124fc6 100644 --- a/docs/source-en/rst_source/overview.rst +++ b/docs/source-en/rst_source/overview.rst @@ -26,6 +26,26 @@ Together, these principles allow RPent to move beyond traditional robot control and establish an agentic infrastructure for the physical world, where intelligence is not only deployed, but continuously built, expanded, and evolved. +Benchmark Results +----------------- + +Compare success rates on LIBERO, LIBERO-PRO, RoboCasa365 Target50, and RoboTwin +C2R. Rankings apply to the methods and evaluation coverage shown; see +:doc:`benchmarks` for detailed results, configurations, and sources. + +.. image:: https://cdn.jsdelivr.net/gh/RLinf/misc@c3b9b5d4ffa360a8324c5b7aa510e1ed0876aa43/rpent/benchmarks/leaderboard-en-light.png + :alt: RPent benchmark results + :class: only-light + :width: 100% + :target: benchmarks.html + +.. image:: https://cdn.jsdelivr.net/gh/RLinf/misc@c3b9b5d4ffa360a8324c5b7aa510e1ed0876aa43/rpent/benchmarks/leaderboard-en-dark.png + :alt: RPent benchmark results + :class: only-dark + :width: 100% + :target: benchmarks.html + + Feature Matrix -------------- diff --git a/docs/source-en/rst_source/usage/libero.rst b/docs/source-en/rst_source/usage/libero.rst index df38cab1b..86101b7c4 100644 --- a/docs/source-en/rst_source/usage/libero.rst +++ b/docs/source-en/rst_source/usage/libero.rst @@ -278,11 +278,16 @@ See :doc:`../development/add_primitive` for the full walkthrough. Reproducing results ------------------- -The following results reproduce -:doc:`Harness VLA <../awesome_works/harnessvla>` on two LIBERO-PRO suites. -On the `reproduce/libero -`_ branch, use -``gpt-5.5`` to reproduce these results: +See :doc:`../benchmarks` for the unified RPent model comparison on LIBERO-PRO +Task/Swap and the corresponding model configurations. + +The :doc:`GPT-6 Astra suite results <../benchmarks>` +cover all eight complete suites and 800 verified episodes: 741 successes, +59 failures, and 92.63% Overall, with Codex / GPT-6 Astra / low / reasoning. + +The following historical reproduction records use the `reproduce/libero +`_ branch with +``gpt-5.5`` and ``xhigh`` reasoning effort: - ``libero_10_task``: 70% (70/100) - ``libero_10_swap``: 55% (55/100) diff --git a/docs/source-zh/index.rst b/docs/source-zh/index.rst index e6f74c37a..9f1c04d57 100644 --- a/docs/source-zh/index.rst +++ b/docs/source-zh/index.rst @@ -66,6 +66,7 @@ 概览 安装 快速开始 + 基准测试结果 .. toctree:: :maxdepth: 2 diff --git a/docs/source-zh/rst_source/benchmarks.rst b/docs/source-zh/rst_source/benchmarks.rst new file mode 100644 index 000000000..336ac3e1c --- /dev/null +++ b/docs/source-zh/rst_source/benchmarks.rst @@ -0,0 +1,24 @@ +:html_theme.sidebar_secondary.remove: + +.. _benchmark-results: +.. _benchmark-leaderboard: +.. _leaderboard: + +RPent 排行榜 +================= + +.. raw:: html + + + + + +
+
+

LIBERO-PRO

方法 / 模型成功率
Codex / GPT-6 Astra / low / reasoning92.63%
Claude Code / Opus-4.7 / max.reasoning82.4%
RPent Flash Mode72.63%
Codex / GPT-5.5 / xhigh / reasoning72.1%
ASPIRE61.36%
π_RLinf50.0%
π0.511.0%
AtomVLA6.3%
X-VLA3.8%
MolmoAct1.5%
π00.3%

GPT-6 Astra: 92.63% (741/800). RPent Flash Mode / Molmo2-8B: 72.63% (581/800).

完整方法与分项成绩

方法 / 模型总体Spatial TaskSpatial SwapObject TaskObject SwapGoal TaskGoal SwapLong TaskLong Swap
GPT-6 Astra92.63%100%98%100%99%88%99%85%72%
Opus-4.7 / max.reasoning82.4%
RPent Flash Mode / Molmo2-8B72.63%79.00%70.00%86.00%93.00%74.00%65.00%60.00%54.00%
GPT-5.572.1%81.0%69.0%94.0%91.0%75.0%66.0%52.0%49.0%
ASPIRE61.36%60.0%51.0%95.0%98.0%45.0%81.0%38.3%22.6%
π_RLinf50.0%42.0%59.0%71.0%78.0%45.0%42.0%49.0%14.0%
π0.511.0%1.0%20.0%1.0%17.0%2.0%38.0%1.0%8.0%
AtomVLA6.3%1.0%16.0%0.0%10.0%11.0%2.0%9.0%1.0%
X-VLA3.8%0.0%0.0%8.0%2.0%9.0%1.0%10.0%0.0%
MolmoAct1.5%0.0%0.0%0.0%6.0%0.0%0.0%6.0%0.0%
π00.3%0.0%0.0%0.0%2.0%0.0%0.0%0.0%0.0%
Cap-X14.0%12.0%18.0%22.0%17.0%26.0%
RATS31.0%29.0%63.0%61.0%36.0%43.0%

ASPIRE:Long Task 和 Long Swap 使用 LIBERO-90 技能库进行 zero-shot 迁移。

GPT-6 Astra:Long Task/Swap 与其余六套件使用各自探索后冻结的 memory 文件快照,评测期间不更新。Overall 合并两个不重叠批次:Long 157/200,加上其余套件 584/600,得到 741/800(92.63%);并非全部回合共享同一份 memory 快照。

+

LIBERO

方法 / 模型成功率
AtomVLA97.0%
Claude Code / Opus-4.7 / max.reasoning96.0%
π_RLinf95.3%
π094.2%
NORA79.5%
OpenVLA76.5%
+

RoboCasa365 · Target50

方法 / 模型成功率
Codex / GPT-6 Astra / low / reasoning59.20%
Xiaomi-Robotics-157.4%
Codex / GPT-5.5 / xhigh / reasoning57.1%
Claude Code / Opus-4.7 / max.reasoning48.6%
WorldDreamer35.3%
RLDX-130.0%
π0.516.9%
π014.8%

RoboCasa365 · Target50 · 完整方法与分项成绩

方法 / 模型总体Atomic-SeenComposite-SeenComposite-Unseen
GPT-6 Astra59.20%87.78%43.75%42.50%
Xiaomi-Robotics-157.4%80.2%57.1%32.1%
GPT-5.557.1%92.0%61.0%13.8%
Opus-4.7 / max.reasoning48.6%79.4%47.5%15.0%
WorldDreamer35.3%66.3%26.7%9.0%
RLDX-130.0%60.0%21.3%5.0%
π0.516.9%39.6%7.1%1.2%
π014.8%34.6%6.1%1.1%
+

RoboTwin

方法 / 模型成功率
Codex / GPT-5.5 / xhigh / reasoning62.4%
Claude Code / Opus-4.7 / max.reasoning58.4%
LingBot-VLA50.4%
π0.547.9%
GR00T-N1.720.7%
StarVLA10.6%
+
+
diff --git a/docs/source-zh/rst_source/overview.rst b/docs/source-zh/rst_source/overview.rst index 89946e887..af8c0749d 100644 --- a/docs/source-zh/rst_source/overview.rst +++ b/docs/source-zh/rst_source/overview.rst @@ -19,6 +19,25 @@ RPent 建立在三条核心设计原则之上: **服务化、标准化、可组 智能体基础设施 (agentic infrastructure for the physical world) —— 在这里, 智能不只是被部署, 而是被持续构建、扩展与演进。 +基准测试结果 +------------ + +对比 LIBERO、LIBERO-PRO、RoboCasa365 Target50 和 RoboTwin C2R 上的成功率。 +排名仅限图中方法及评测范围;完整结果、模型配置和来源见 :doc:`benchmarks`。 + +.. image:: https://cdn.jsdelivr.net/gh/RLinf/misc@c3b9b5d4ffa360a8324c5b7aa510e1ed0876aa43/rpent/benchmarks/leaderboard-zh-light.png + :alt: RPent benchmark results + :class: only-light + :width: 100% + :target: benchmarks.html + +.. image:: https://cdn.jsdelivr.net/gh/RLinf/misc@c3b9b5d4ffa360a8324c5b7aa510e1ed0876aa43/rpent/benchmarks/leaderboard-zh-dark.png + :alt: RPent benchmark results + :class: only-dark + :width: 100% + :target: benchmarks.html + + 功能矩阵 -------- diff --git a/docs/source-zh/rst_source/usage/libero.rst b/docs/source-zh/rst_source/usage/libero.rst index 1ee4003d6..172395647 100644 --- a/docs/source-zh/rst_source/usage/libero.rst +++ b/docs/source-zh/rst_source/usage/libero.rst @@ -254,8 +254,13 @@ Dashboard 支持 ``api``、``claude_code`` 和 ``codex`` planner。 结果复现 -------- -以下是在两个 LIBERO-PRO 套件上复现 -:doc:`Harness VLA <../awesome_works/harnessvla>` 得到的结果。实验使用 +RPent 在 LIBERO-PRO Task/Swap 上的统一模型对比及对应配置见 :doc:`../benchmarks`。 + +:doc:`GPT-6 Astra 套件汇总 <../benchmarks>` +记录全部八个完整套件及 800 个已核验回合:741 成功、59 失败,Overall 92.63%, +配置为 Codex / GPT-6 Astra / low / reasoning。 + +以下保留历史复现记录,实验使用 `reproduce/libero `_ 分支和 ``gpt-5.5`` 模型: