Making Progress Measurable

Evaluation and Benchmarking of Driving Policies

Kashyap Chitta

ELLIS Institute Tübingen and KE:SAI

2026-09-16

Where we are and how we got here

Autonomous driving in the 1990s

Ernst Dickmanns and team, Universität der Bundeswehr München (Siedersberger 2003)

1,758 km

Munich–Odense–Munich

95%

of the distance autonomous

1995

On public highways

Autonomous driving in the 2020s

Tesla FSD 14.3

NVIDIA DRIVE AV

Horizon HSD 2.0

Xpeng VLA 2.0

WeRide WRD 3.0

Nio NWD 2.0

The automotive industry spent >$100B to automate 0.00006% of traffic.

Open source investment accelerated digital AI

The open-source wave in autonomous driving is here

Returns on open-source investment

From open source software to open source strategy (Gurley 2026)

Bill Gurley's overview of open AV stakeholders, the risk of dependence on proprietary stacks, and a proposed global open-source consortium.

Activating this trillion-dollar coalition through open-source will benefit all members.

Benchmarks catalyze progress

Linux: comparisons led to funding (Kegel 2003; VeriTest 2003)

Mindcraft (1999)

Windows NT 4.0 / IIS vs. Red Hat 5.2 / Apache

2.5× Windows advantage.

VeriTest (2003)

Windows Server 2003 vs. Red Hat 2.1

1.66× Windows advantage.

Open code attracts sustained investment when businesses depend on it.

Accelerating open research with the common task framework

The story of speech transcription

1970s: the demo culture in speech technology

Every group chooses its own test, and every approach looks promising

  • Elaborate speech systems that consumed resources without yielding clear knowledge (Pierce 1969).
  • Pushback from executives at Bell Labs and US government on potential value.
  • Zero US funding for speech or translation research from 1975-1985.

1985–1991: speech adopts a common task framework

Same recordings, same scoring rules, comparable results (Pallett 1985; Wayne 1991)

  • 1985: Pallett advocates the common task framework.
  • 1991: NIST compares systems under this framework on shared Resource Management (RM) and Air Travel Information System (ATIS) corpora.

Share

Common dev split of speech recordings and transcripts

Test

Held-out recordings, scored by pre-defined metric

Compare

Anything goes, best method wins

Not everyone liked it

Reactions to the common task framework (Liberman 2015)

Many were skeptical

“You can’t turn water into gasoline, no matter what you measure.”

Others were disgruntled

“It’s like being in first grade again — you’re told exactly what to do, and then you’re tested over and over.”

But it worked!

Shared tests made progress visible (National Institute of Standards and Technology 2009)

NIST speech-to-text benchmark history, May 2009. Word error rate on a logarithmic axis versus year, with separate series for read, air-travel, broadcast, conversational, and meeting speech. Many series show declining error, while harder tasks retain higher error rates.

The common task framework has cemented itself

In speech, vision (Liu et al. 2019), language, and some sciences

  • 1990s: no papers had benchmarks
  • 2010s: no papers without benchmarks

Figure 3 from Liu et al.: mean average precision of winning object detection systems on PASCAL VOC 2007–2012 and ILSVRC 2013–2017, showing substantial improvements with deep learning.

The battle for common task frameworks in driving

CARLA releases and challenges

The most impactful simulator for embodied AI

CARLA GitHub star history from 2017 to September 2026, rising to about 14,400 stars. Vertical version labels mark all 27 three-part releases on the blue curve, from 0.5.4 to 0.9.16, including 0.10.0. Flame-colored markers identify Challenge 2019–2024 at their host conferences’ start dates.

Evaluating entire driving stacks is hard

CARLA did not define standard local datasets or benchmarks (Jaeger et al. 2024)

  • Matching permitted input spaces, conditioning, training towns, evaluation routes, safety-critical scenarios, and traffic density was non-trivial.
  • Training on validation towns or disabling scenarios changes the task.
  • Some papers introduce easier benchmarks under an existing benchmark’s name, making their methods appear stronger than the results justify.
  • Papers claimed state of the art while trailing methods published years earlier.
  • Errors propagated when authors copied tables from other papers.

We need an independent evaluator to run standardized local evaluations too.

Verifiable != reproducible

An independent server verifies the result, but cannot verify the paper’s explanation.

  • Three leading Leaderboard 1.0 methods reproduced substantially below their reported scores, despite public code and models matching the papers.
  • Their released system differed from the submitted system.

We need to enforce or encourage reproducibility.

It then became prohibitively expensive

Closed-loop outcomes vary with both simulation and model training.

  • End-to-end methods can also have substantial training variance. Repeating simulation with one trained model does not measure that uncertainty.

3

Independently trained models

× 3

Evaluation repeats per model

= 9

Runs averaged

We need evaluation cheap enough to measure both training and simulation variance.

The breakdown of the common task framework

nuScenes planning and the return of mainstream open-loop evaluation

Slide 6 from the CVPR 2024 NAVSIM talk: Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving? Trajectory heatmap and typical nuScenes scene.

Slide 7 from the CVPR 2024 NAVSIM talk: Ego-MLP.

We need to define tasks that stop rewarding shortcut learning.

Takeaways

What a common task framework for driving needs

  • Task design: stop rewarding shortcut learning.
  • Affordability: cheap enough to measure training and simulation variance.
  • Reproducibility: enforce or encourage reproducible results.
  • Independent evaluation: standardized evaluations on an independent server.

Lightweight simulation with NAVSIM

Distance from a human is an unreliable target

Evaluating safety, comfort, and progress (Dauner et al. 2024)

Improving the test distribution

Filtering results in diverse scenes

Slide 27 from the CVPR 2024 NAVSIM talk: unfiltered recordings are mostly static or straight driving scenes.

Slide 28 from the CVPR 2024 NAVSIM talk: filtered data results in more diverse and challenging scenes.

Shortcut learning from ego history ineffective.

Evaluating several possible futures for each test sample

3D Gaussian Splatting to measure recovery capability (Cao et al. 2025).

Figure 2, scene 1: real and synthetic front-camera views linked to vehicle poses on a curved road. Figure 2, scene 2: real and synthetic front-camera views with their corresponding vehicle poses. Figure 2, scene 3: alternative future viewpoints and their positions in the road layout. Figure 2, scene 4: alternative future viewpoints and their positions in the road layout.

Pre-rendering and deterministic simulation improves affordability.

Weighting observations by proximity

Avoiding unsolvable states to prevent their penalties

Pseudo-simulation uses pre-rendered observations around possible future vehicle states

Pseudo-simulation predicts closed-loop performance

Displacement-based metrics severely underestimate rule-based planners

Figure 3: scatter plots compare ADE, nuPlan open-loop score, and pseudo-simulation EPDMS with closed-loop scores. EPDMS shows substantially stronger alignment across rule-based and learned planners.

Live NAVSIM leaderboards

Closed-loop simulation at scale with AlpaSim

The AlpaSim challenge team

Kashyap Chitta
Kashyap Chitta
Michael Watson
Michael Watson
Long Nguyen
Long Nguyen
William Lew
William Lew
Frieda Rong
Frieda Rong
Caojun Wang
Caojun Wang
Haochen Tian
Haochen Tian
Yihang Qiu
Yihang Qiu
Boris Ivanovic
Boris Ivanovic
Maximilian Igl
Maximilian Igl
Yiyi Liao
Yiyi Liao
Peter Karkus
Peter Karkus
Andrei Bursuc
Andrei Bursuc
and many more

Two tracks with a single interface

Physical AI AV (1700 train hours)

nuPlan (120 train hours)

One devkit for any training dataset

QR code linking to py123d_garage on GitHub

A common task framework with AlpaSim

Realistic constraints: 16 GiB GPU memory and 10 Hz target planning frequency

Affordability

  • Pipeline parallelism
  • Constant engineering gains
  • Organizer compute for challenge

Independent evaluation

  • Thousands of private held-out scenes across ten countries
  • Smaller, standardized local splits

Reproducibility

  • Tech reports checked against containers
  • Periodic cleanup for public leaderboard

Live AlpaSim leaderboard

QR code linking to the AlpaSim challenge landing page

Open research in benchmarking driving policies

Standardizing deployment to the real-world

From simulation research vehicle deployment

Side view of the KE:SAI VW ID.7 research vehicle with its roof-mounted sensor rack. Close-up of cameras and LiDAR sensors on the roof rack. Onboard computer and networking equipment installed in the trunk.

Multi-fidelity benchmarking

Multi-domain benchmarking

More diverse benchmarks can have less stable rankings (Zhang and Hardt 2024)

Learning to weight tasks

Item Response Theory: estimate ability, difficulty, and uncertainty together (Birnbaum 1968)

Context-aware evaluation

Frontier models can decide which rules matter (Sun et al. 2026)

DriveJudge uses video, a bird's-eye view, and rule-based metrics to select relevant rules and evaluate a trajectory. When nudging around a parked bus, strict lane keeping is not relevant; the remaining metrics contribute to the aggregate score.

References

Birnbaum, A. (1968). Some latent trait models and their use in inferring an examinee’s ability. In F. M. Lord & M. R. Novick (Eds.), Statistical theories of mental test scores (pp. 397–479). Addison-Wesley.
Cao, W., Hallgarten, M., Li, T., Dauner, D., Gu, X., Wang, C., et al. (2025). Pseudo-simulation for autonomous driving. In Conference on robot learning (CoRL).
Dauner, D., Hallgarten, M., Li, T., Weng, X., Huang, Z., Yang, Z., et al. (2024). NAVSIM: Data-driven non-reactive autonomous vehicle simulation and benchmarking. In Advances in neural information processing systems (NeurIPS).
Gurley, B. (2026, May 9). From open source software to open source strategy. https://p3institute.substack.com/p/from-open-source-software-to-open
Jaeger, B., Chitta, K., Dauner, D., Renz, K., & Geiger, A. (2024, December 16). Common mistakes in benchmarking autonomous driving. https://github.com/autonomousvision/carla_garage/blob/leaderboard_2/docs/common_mistakes_in_benchmarking_ad.md#carla-benchmarks
Kegel, D. (2003). On Mindcraft’s april 1999 benchmark. https://www.kegel.com/mindcraft_redux.html
Liberman, M. (2015). Reproducible computational experiments. CATS Reproducibility Workshop presentation. https://languagelog.ldc.upenn.edu/myl/LibermanCATS02262015.pdf
Liu, L., Ouyang, W., Wang, X., Fieguth, P., Chen, J., Liu, X., & Pietikäinen, M. (2019). Deep learning for generic object detection: A survey. arXiv preprint arXiv:1809.02165v4. https://arxiv.org/html/1809.02165v4
National Institute of Standards and Technology. (2009). NIST STT benchmark test history—may 2009. https://www.nist.gov/sites/default/files/images/2016/12/22/history-1slide-may09-single.png
Pallett, D. S. (1985). Performance assessment of automatic speech recognizers. Journal of Research of the National Bureau of Standards, 90(5), 371–387. https://nvlpubs.nist.gov/nistpubs/jres/090/jresv90n5p371_A1b.pdf
Pierce, J. R. (1969). Whither speech recognition? The Journal of the Acoustical Society of America, 46(4B), 1049–1051. https://doi.org/10.1121/1.1911801
Red Hat. (2012). Red hat reports fourth quarter and fiscal year 2012 results. https://www.redhat.com/en/about/press-releases/red-hat-reports-fourth-quarter-and-fiscal-year-2012-results
Siedersberger, K.-H. (2003). Komponenten zur automatischen fahrzeugführung in sehenden (semi-)autonomen fahrzeugen (PhD thesis). Universität der Bundeswehr München. Retrieved from https://athene-forschung.unibw.de/doc/85329/85329.pdf
Sun, X., Xie, K., Schmalfuss, J., Paschalidou, D., Zhang, X., Fidler, S., et al. (2026). DriveJudge: Rethinking autonomous driving evaluation with vision-language models. https://arxiv.org/abs/2606.17362v1
The Linux Foundation. (2009). Linux foundation expands individual membership program with new benefits and linux.com email address. https://www.linuxfoundation.org/press/press-release/linux-foundation-expands-individual-membership-program-with-new-benefits-and-linux-com-email-address
VeriTest. (2003). Microsoft windows server 2003 vs. Linux competitive file server performance comparison. VeriTest. https://tech-insider.org/windows/research/acrobat/0304-a.pdf
Wayne, C. L. (1991). A snapshot of two DARPA speech and natural language programs. In Speech and natural language: Proceedings of a workshop held at pacific grove, california, february 19–22, 1991. https://aclanthology.org/H91-1078/
Zhang, G., & Hardt, M. (2024). Inherent trade-offs between diversity and stability in multi-task benchmarks. In International conference on machine learning (ICML). https://arxiv.org/abs/2405.01719