Preprint · 2026

World4Scorer

Outcome-Grounded World Modeling for Autonomous Driving

Jieyuan Pei1,2,3*‡Meiyi Lu4*Sining Ang5Yubo Zhao6Zhangyi Hu3Mingwei Xu7Haokai Ding8Wei Li9Zihan You1,10Jianwei Zheng9Li Yu11Yifeng Pan11Ji Tao11Rongjunchen Zhang2Yan Wang1†

1Institute for AI Industry Research (AIR), Tsinghua University2HiThink Research3The Hong Kong University of Science and Technology (Guangzhou)4Zhejiang University5University of Science and Technology of China6SMBU7University of Washington8Mohamed bin Zayed University of Artificial Intelligence9Zhejiang University of Technology10Southeast University11Changan Automobile

*Equal contribution · †Corresponding author · ‡Work done during an internship

Overview

The scorer is a world model

A generate-and-select planner proposes many trajectories and must pick one it has never seen executed. World4Scorer predicts a state for every candidate, trains those states with simulator outcomes, and anchors the shared predictor with the one future the driving log actually recorded.

Driving logs record the future of only the executed trajectory, so matching the logged future can leave predictions for the alternatives unconstrained. A simulator, in contrast, can label the outcome of every candidate.

World4Scorer builds the scorer as a trajectory-conditioned JEPA-style predictor. Outcome labels supervise the predicted state of every candidate, and the observed future anchors the shared predictor during training only.

Read the full abstract →

Autonomous driving requires choosing a safe and efficient plan as surrounding traffic evolves. Generate-and-select planners propose multiple trajectories and score them for execution, and they have outperformed representative direct-prediction baselines on NAVSIM. Their scorer must compare plans that were never executed. Driving logs record the future of only the executed trajectory, so matching the logged future can leave predictions for the alternatives unconstrained; a simulator, in contrast, can label the outcome of every candidate. We introduce World4Scorer, which builds the scorer as a trajectory-conditioned JEPA-style predictor: it predicts a state for each candidate and reads the candidate's scores from that state. Simulator outcome labels supervise the states of all candidates, and the observed future of the executed trajectory anchors the predictor to real scene evolution. Because one predictor produces every candidate's state, the anchor can constrain shared parameters used to score unexecuted plans, while the future itself is needed only during training. Generated candidates mostly score well, so a scene-matched bank adds low-scoring plans to the outcome supervision; framewise choices can conflict, so inertial re-ranking keeps consecutive selections consistent. World4Scorer achieves state-of-the-art NAVSIM-v2 performance and a strong adapted-system result on closed-loop Bench2Drive. With the LeWM world model and planning budget fixed, outcome-based scoring also improves manipulation planning on the OGBench-Cube benchmark.

Outcome labels reach every candidate

Collision, drivable area, progress, time-to-collision and comfort labels for all 64 generated and 16 bank plans.

The observed future anchors the predictor

One logged future supervises the executed query; shared parameters carry the constraint to the other candidates.

No future needed at test time

At deployment the predictor sees only current observations and candidate trajectories; inertial re-ranking adds no learned parameters.

Live demo

Watch it score 64 plans every half second

A replay of World4Scorer on recorded NAVSIM drives. Every 0.5 s the generator proposes 64 trajectories; the predicted score colours each one, and the purple plan is the scorer's choice. Drag the timeline, change the speed, or switch layers.

front camera0 km/h
bird's-eye view · heading up
0.0 / 0.0 s
predicted score, low → high World4Scorer plan human driver
Method

From scoring plans to predicting their outcomes

Three ways to train the scorer of a generate-and-select planner, and the architecture that combines outcome labels with future prediction.

Generate and select
outcome supervision future-prediction supervision simulator label observed future unobserved future training only

World4Scorer architecture: DINOv2 scene tokens, a trajectory generator, a shared latent world model mapping each candidate to a state, score heads, a visual readout, and inertial re-ranking.

Architecture. (a) The generator proposes 64 candidates; one predictor maps each to a state, score heads read its outcomes, and inertial re-ranking checks continuity with the previous plan. (b) Joint training: the visual readout of the executed query predicts the frozen DINOv2 feature two seconds ahead, while outcome labels supervise all generated and 16 bank candidates.

01

Outcome-grounded states

Every candidate gets a predicted state; simulator outcomes supervise all of them, including plans that were never driven.

02

Complementary supervision

An observed-future target improves planning where a current-frame target does not, and correctly paired outcome labels are essential.

03

Inertial re-ranking

A training-free check of each candidate against the previous plan keeps consecutive choices consistent: +1.6 EPDMS.

Theory

Why one observed future helps every candidate

Several scorers can fit all the outcome labels and still disagree about which candidate to pick. Around a fitted model, we ask how far a candidate comparison can move within the region the training losses allow.

  1. Score-gap sensitivity. For candidates i and j, let . Its radius is the largest first-order change of the score gap over the trust region set by the loss curvature .
  2. The future can only tighten it. Adding the future loss adds curvature, so the radius never grows, and it shrinks strictly when the executed query's readout moves along the comparison direction.
  3. Stable decisions. If a candidate leads by more than the radius, the linearized ranking cannot flip anywhere in that region.

strictly if and only if , where is the Jacobian of the normalized readout.

A local, first-order result under a generalized Gauss–Newton surrogate of the outcome and future losses (Proposition 1 and its proof in the appendix of the paper).

Two score surfaces fit the outcome labels but favour different candidates; future constraints remove that disagreement.

Complementary supervision. Two score surfaces fit the outcome labels but favour different candidates. Future constraints remove that disagreement while allowing different predictions elsewhere.

Results

State of the art on NAVSIM-v2, strong in closed loop

All numbers as reported in the paper. Our NAVSIM evaluations use the latest official devkits.

PDMS (NAVSIM-v1) and EPDMS (NAVSIM-v2) on navtesthover a point for details

NAVSIM navtest, 12,146 scenes. Input C: cameras, L: LiDAR. Click a column header to sort. As in Table 1 of the paper.

Bench2Drive Driving Score, 220 closed-loop routes

World4Scorer success rate 46.36%. Our adaptation adds a route-point input and uses the simulator ego state; reference results as published.

OGBench-Cube success rate (%)

GCIQL, GCIVL and PLDM as reported by LeWM. LeWM and World4Scorer are our runs with the same frozen world model and CEM planning budget, averaged over 11 seeds of 50 episodes; paired gain +4.73 points, 95% CI [+2.18, +7.27].

Citation

BibTeX

@article{pei2026world4scorer,
  title   = {World4Scorer: Outcome-Grounded World Modeling for Autonomous Driving},
  author  = {Pei, Jieyuan and Lu, Meiyi and Ang, Sining and Zhao, Yubo and Hu, Zhangyi and Xu, Mingwei and Ding, Haokai and Li, Wei and You, Zihan and Zheng, Jianwei and Yu, Li and Pan, Yifeng and Tao, Ji and Zhang, Rongjunchen and Wang, Yan},
  journal = {arXiv preprint arXiv:2609.36438},
  year    = {2026}
}