Skip to content

Commit 3eded7e

Browse files
committed
Update README
1 parent 46fbc48 commit 3eded7e

1 file changed

Lines changed: 20 additions & 3 deletions

File tree

‎README.md‎

Lines changed: 20 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -55,6 +55,7 @@
5555

5656

5757
## News
58+
- **[2026/09/15]** Stable checkpoints released.
5859
- **[2026/06/19]** Paper released on arXiv. See [World Engine: Towards the Era of Post-Training for Autonomous Driving](https://arxiv.org/abs/2606.19836).
5960
- **[2026/04/09]** Official dataset released. See [OpenDriveLab/WorldEngine](https://huggingface.co/datasets/OpenDriveLab/WorldEngine) or [OpenDriveLab/WorldEngine (ModelScope)](https://www.modelscope.cn/datasets/OpenDriveLab/WorldEngine)
6061
- **[2026/04/10]** Official code repository established.
@@ -65,16 +66,17 @@
6566
We compare different post-training paradigms on the nuPlan dataset, evaluating on both open-loop and closed-loop metrics across common and rare driving scenarios.
6667

6768
> **Metric notes:**
68-
> **Early stage**. Stable ckpts and corresponding results coming soon.
6969
> - **Open-loop PDMS** is aligned with [NAVSIM v1.1](https://github.com/autonomousvision/navsim) PDM Score. *Common* denotes the standard `navtest` split; *Rare* denotes the `navtest_failures` subset — failure-prone rare-case scenarios extracted from `navtest`.
70-
> - **Closed-loop Success Rate** is defined as the fraction of simulated driving episodes completed without collision or off-road failure.
70+
> - **Closed-loop Success Rate (SR)** is computed as NC × DAC (no-at-fault-collision score × drivable-area compliance score), reported as a percentage.
7171
> - **Closed-loop Ego Progress (EP)** measures the route progress made by the ego vehicle during **SimEngine closed-loop testing**, reflecting whether the agent makes meaningful forward progress rather than merely avoiding collision or off-road failure.
7272
> - **Closed-loop PDMS*** is the PDM Score obtained via **SimEngine closed-loop testing**, where the planner interacts with reactive agents in simulation under real-time rendering.
7373
>
7474
> **Training notes:**
7575
> - **Rare logs** are failure-prone scenarios automatically extracted from `navtrain` by the pre-trained agent itself (see [Rare Case Extraction](docs/algengine_usage.md#rare-case-extraction)).
7676
> - **Common logs** are the standard cases in `navtrain`.
7777
78+
#### VADv2 (Base Planner)
79+
7880
| Method | Open-loop PDMS ↑ (common) | Open-loop PDMS ↑ (rare) | Closed-loop SR ↑ (rare) | Closed-loop EP ↑ (rare) | Closed-loop PDMS* ↑ (rare) |
7981
|:-------|:-------------------------:|:-----------------------:|:-----------------------:|:-----------------------:|:--------------------------:|
8082
| Base model | 85.64 | 47.14 | 73.66 | 46.71 | 60.98 |
@@ -90,6 +92,21 @@ We compare different post-training paradigms on the nuPlan dataset, evaluating o
9092
- Post-training on **common logs** provides limited long-tail benefit and degrades rare closed-loop performance, reducing SR from **73.66%** to **69.63%** and PDMS$^\ast$ from **60.98** to **60.21**, confirming the importance of long-tail event discovery.
9193
- The full WorldEngine pipeline achieves the best overall rare closed-loop performance, with the highest SR (**88.89%**) and PDMS$^\ast$ (**70.12**). It improves rare closed-loop SR by **+15.23** percentage points and PDMS$^\ast$ by **+9.14** over the base model, while maintaining strong common open-loop performance.
9294

95+
#### HydraMDP
96+
97+
The base model and WorldEngine rows match **Table S1** of the paper. Additional ablations follow the selected HydraMDP experiment records: pure IL fine-tuning for the supervised row, and reward shaping enabled, RL fine-tuning enabled, PG = 0.01, entropy = 0 for the RL rows. Closed-loop values use the **reactive** evaluation results.
98+
99+
| Method | Open-loop PDMS ↑ (common) | Open-loop PDMS ↑ (rare) | Closed-loop SR ↑ (rare) | Closed-loop EP ↑ (rare) | Closed-loop PDMS* ↑ (rare) |
100+
|:-------|:-------------------------:|:-----------------------:|:-----------------------:|:-----------------------:|:--------------------------:|
101+
| Base model | 93.86 | 69.91 | 75.15 | 62.98 | 68.69 |
102+
| Supervised fine-tuning on rare logs | 88.49 | 57.43 | 70.77 | 52.81 | 62.74 |
103+
| Post-training on common logs | 93.90 | 69.64 | 76.61 | 64.80 | 68.70 |
104+
| Post-training on rare synthetic replays | 93.87 | 69.91 | 74.26 | 62.01 | 67.68 |
105+
| Post-training on rare rollouts w/o Behaviour WM | 93.78 | 71.59 | 78.57 | 67.60 | 72.06 |
106+
| **Post-training with WorldEngine** | **93.89** | **72.28** | **81.63** | **67.95** | **74.49** |
107+
108+
WorldEngine improves HydraMDP's rare closed-loop SR by **+6.48** percentage points, EP by **+4.97**, and PDMS$^\ast$ by **+5.80**, matching the gains reported in Table S1.
109+
93110
### Qualitative Results — Closed-Loop Simulation on nuPlan
94111

95112
Each pair shows the **Base model** vs **WorldEngine post-trained model** on the same rare-case scenario. Left: front-camera rendering; Right: BEV trajectory visualization.
@@ -138,8 +155,8 @@ WorldEngine consists of two tightly coupled subsystems:
138155
- [x] Hugging Face / ModelScope dataset
139156
- [x] Open-source release (code, data, early pre-trained models)
140157
- [x] arXiv preprint
158+
- [x] Stable pre-trained models
141159
- [ ] Behavior World Model integration
142-
- [ ] Stable pre-trained models
143160

144161

145162
## Getting Started

0 commit comments

Comments
 (0)