Birnbaum, A. (1968). Some latent trait models and their use in inferring an examinee’s ability. In F. M. Lord & M. R. Novick (Eds.), Statistical theories of mental test scores (pp. 397–479). Addison-Wesley.
Cao, W., Hallgarten, M., Li, T., Dauner, D., Gu, X., Wang, C., et al. (2025). Pseudo-simulation for autonomous driving. In Conference on robot learning (CoRL).
Dauner, D., Hallgarten, M., Li, T., Weng, X., Huang, Z., Yang, Z., et al. (2024). NAVSIM: Data-driven non-reactive autonomous vehicle simulation and benchmarking. In Advances in neural information processing systems (NeurIPS).
Gurley, B. (2026, May 9). From open source software to open source strategy.
https://p3institute.substack.com/p/from-open-source-software-to-open
Jaeger, B., Chitta, K., Dauner, D., Renz, K., & Geiger, A. (2024, December 16). Common mistakes in benchmarking autonomous driving.
https://github.com/autonomousvision/carla_garage/blob/leaderboard_2/docs/common_mistakes_in_benchmarking_ad.md#carla-benchmarks
Kegel, D. (2003). On
Mindcraft’s april 1999 benchmark.
https://www.kegel.com/mindcraft_redux.html
Liberman, M. (2015). Reproducible computational experiments. CATS Reproducibility Workshop presentation.
https://languagelog.ldc.upenn.edu/myl/LibermanCATS02262015.pdf
Liu, L., Ouyang, W., Wang, X., Fieguth, P., Chen, J., Liu, X., & Pietikäinen, M. (2019). Deep learning for generic object detection: A survey.
arXiv preprint arXiv:1809.02165v4.
https://arxiv.org/html/1809.02165v4
National Institute of Standards and Technology. (2009).
NIST STT benchmark test history—may 2009.
https://www.nist.gov/sites/default/files/images/2016/12/22/history-1slide-may09-single.png
Pallett, D. S. (1985). Performance assessment of automatic speech recognizers.
Journal of Research of the National Bureau of Standards,
90(5), 371–387.
https://nvlpubs.nist.gov/nistpubs/jres/090/jresv90n5p371_A1b.pdf
Pierce, J. R. (1969). Whither speech recognition?
The Journal of the Acoustical Society of America,
46(4B), 1049–1051.
https://doi.org/10.1121/1.1911801
Siedersberger, K.-H. (2003).
Komponenten zur automatischen fahrzeugführung in sehenden (semi-)autonomen fahrzeugen (PhD thesis). Universität der Bundeswehr München. Retrieved from
https://athene-forschung.unibw.de/doc/85329/85329.pdf
Sun, X., Xie, K., Schmalfuss, J., Paschalidou, D., Zhang, X., Fidler, S., et al. (2026).
DriveJudge: Rethinking autonomous driving evaluation with vision-language models.
https://arxiv.org/abs/2606.17362v1
VeriTest. (2003).
Microsoft windows server 2003 vs. Linux competitive file server performance comparison. VeriTest.
https://tech-insider.org/windows/research/acrobat/0304-a.pdf
Wayne, C. L. (1991). A snapshot of two
DARPA speech and natural language programs. In
Speech and natural language: Proceedings of a workshop held at pacific grove, california, february 19–22, 1991.
https://aclanthology.org/H91-1078/
Zhang, G., & Hardt, M. (2024). Inherent trade-offs between diversity and stability in multi-task benchmarks. In
International conference on machine learning (ICML).
https://arxiv.org/abs/2405.01719