Guide
Bibliography
109 sources. Preprints, reports, working papers, and books are labeled. Everything else is peer-reviewed or a classic in its field.
- Adcock, R., & Collier, D. (2001). Measurement validity: A shared standard for qualitative and quantitative research. American Political Science Review, 95(3), 529-546.Cited in: The Construct Gap, Overview
- Amershi, S., et al. (2019). Guidelines for human-AI interaction. Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI 2019).Cited in: The Team Is the System
- Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., & Mané, D. (2016). Concrete problems in AI safety. arXiv:1606.06565. (Preprint)Cited in: Goodhart’s Law
- Aroyo, L., & Welty, C. (2015). Truth is a lie: Crowd truth and the seven myths of human annotation. AI Magazine, 36(1), 15-24.Cited in: The Gold Standard Myth
- Bansal, G., Wu, T., Zhou, J., Fok, R., Nushi, B., Kamar, E., Ribeiro, M. T., & Weld, D. (2021). Does the whole exceed its parts? The effect of AI explanations on complementary team performance. Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI 2021).Cited in: The Team Is the System, Overview, Glossary
- Bean, A. M., Kearns, R. O., Romanou, A., Hafner, F. S., Mayne, H., et al. (2025). Measuring what matters: Construct validity in large language model benchmarks. Advances in Neural Information Processing Systems (NeurIPS 2025), Datasets and Benchmarks Track.Cited in: The Construct Gap
- Beede, E., Baylor, E., Hersch, F., Iurchenko, A., Wilcox, L., Ruamviboonsuk, P., & Vardoulakis, L. M. (2020). A human-centered evaluation of a deep learning system deployed in clinics for the detection of diabetic retinopathy. Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI 2020).Cited in: The Lab-to-Field Gap, Overview
- Bengio, Y., et al. (2026). International AI Safety Report 2026. International AI Safety Report; arXiv:2602.21012. (Report)Cited in: Evaluation Awareness, The Lab-to-Field Gap, Data Contamination, Goodhart’s Law, Overview
- Beyer, B., Jones, C., Petoff, J., & Murphy, N. R. (Eds.) (2016). Site reliability engineering: How Google runs production systems. O’Reilly Media. (Book)Cited in: Being Pragmatic, Glossary
- Biderman, S., et al. (2024). Lessons from the trenches on reproducible evaluation of language models. arXiv:2405.14782. (Preprint)Cited in: The Reproducibility Rule, Prompt Sensitivity
- Bouthillier, X., et al. (2021). Accounting for variance in machine learning benchmarks. Proceedings of Machine Learning and Systems (MLSys 2021).Cited in: No Error Bars, No Result
- Bowman, S. R., & Dahl, G. E. (2021). What will it take to fix benchmarking in natural language understanding? Proceedings of NAACL-HLT 2021.Cited in: Benchmark Saturation
- Breck, E., Cai, S., Nielsen, E., Salib, M., & Sculley, D. (2017). The ML test score: A rubric for ML production readiness and technical debt reduction. Proceedings of the IEEE International Conference on Big Data.Cited in: Being Pragmatic
- Buolamwini, J., & Gebru, T. (2018). Gender shades: Intersectional accuracy disparities in commercial gender classification. Proceedings of the Conference on Fairness, Accountability and Transparency (FAT* 2018), PMLR 81, 77-91.Cited in: The Averaging Trap
- Burnell, R., et al. (2023). Rethink reporting of evaluation results in AI. Science, 380(6641), 136-138.Cited in: The Averaging Trap, Overview
- Buçinca, Z., Malaya, M. B., & Gajos, K. Z. (2021). To trust or to think: Cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW1), Article 188.Cited in: The Team Is the System
- Campbell, D. T. (1979). Assessing the impact of planned social change. Evaluation and Program Planning, 2(1), 67-90.Cited in: Campbell’s Law
- Card, D., Henderson, P., Khandelwal, U., Jia, R., Mahowald, K., & Jurafsky, D. (2020). With little power comes great responsibility. Proceedings of EMNLP 2020.Cited in: No Error Bars, No Result
- Chang, Y., et al. (2024). A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology, 15(3).Cited in: Overview
- Chiang, W.-L., et al. (2024). Chatbot Arena: An open platform for evaluating LLMs by human preference. Proceedings of the International Conference on Machine Learning (ICML 2024).Cited in: Overview, Glossary
- Clark, E., August, T., Serrano, S., Haduong, N., Gururangan, S., & Smith, N. A. (2021). All that’s ‘human’ is not gold: Evaluating human evaluation of generated text. Proceedings of ACL-IJCNLP 2021.Cited in: The Gold Standard Myth, Overview
- Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281-302.Cited in: The Construct Gap, Glossary
- Dehghani, M., Tay, Y., Gritsenko, A. A., et al. (2021). The benchmark lottery. arXiv:2107.07002. (Preprint)Cited in: The Benchmark Lottery, Campbell’s Law
- Dell’Acqua, F., McFowland III, E., Mollick, E., Lifshitz, H., Kellogg, K., Rajendran, S., Krayer, L., Candelon, F., & Lakhani, K. R. (2026). Navigating the jagged technological frontier: Field experimental evidence of the effects of artificial intelligence on knowledge worker productivity and quality. Organization Science, 37(2).Cited in: The Jagged Frontier
- Dell’Acqua, F., McFowland III, E., Mollick, E., Lifshitz-Assaf, H., Kellogg, K., Rajendran, S., Krayer, L., Candelon, F., & Lakhani, K. R. (2023). Navigating the jagged technological frontier: Field experimental evidence of the effects of AI on knowledge worker productivity and quality. Harvard Business School Working Paper 24-013. (Working paper)Cited in: The Jagged Frontier
- Dijkstra, E. W. (1970). Notes on structured programming (EWD249). Technological University Eindhoven (first circulated 1969; second edition April 1970).Cited in: Presence, Not Absence
- Dror, R., et al. (2018). The hitchhiker’s guide to testing statistical significance in natural language processing. Proceedings of ACL 2018.Cited in: No Error Bars, No Result
- Dwork, C., Feldman, V., Hardt, M., Pitassi, T., Reingold, O., & Roth, A. (2015). The reusable holdout: Preserving validity in adaptive data analysis. Science, 349(6248), 636-638.Cited in: Adaptive Overfitting, Glossary
- Edmondson, A. (1999). Psychological safety and learning behavior in work teams. Administrative Science Quarterly, 44(2), 350-383.Cited in: Being Pragmatic
- Eriksson, M., et al. (2025). Can we trust AI benchmarks? An interdisciplinary review of current issues in AI evaluation. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society (AIES 2025).Cited in: The Benchmark Lottery, Overview
- Ethayarajh, K., & Jurafsky, D. (2020). Utility is in the eye of the user: A critique of NLP leaderboards. Proceedings of EMNLP 2020.Cited in: The Benchmark Lottery, The Cost Frontier, Campbell’s Law
- European Union (2024). Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union. (Report)Cited in: Being Pragmatic
- Ferrari Dacrema, M., Cremonesi, P., & Jannach, D. (2019). Are we really making much progress? A worrying analysis of recent neural recommendation approaches. Proceedings of the ACM Conference on Recommender Systems (RecSys 2019).Cited in: The Baseline Rule
- Ganguli, D., et al. (2022). Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned. arXiv:2209.07858. (Preprint)Cited in: Presence, Not Absence, Overview
- Gao, L., Schulman, J., & Hilton, J. (2023). Scaling laws for reward model overoptimization. Proceedings of the International Conference on Machine Learning (ICML 2023).Cited in: Goodhart’s Law
- Gebru, T., et al. (2021). Datasheets for datasets. Communications of the ACM, 64(12), 86-92.Cited in: The Reproducibility Rule, Glossary
- Geirhos, R., Jacobsen, J.-H., Michaelis, C., Zemel, R., Brendel, W., Bethge, M., & Wichmann, F. A. (2020). Shortcut learning in deep neural networks. Nature Machine Intelligence, 2, 665-673.Cited in: The Clever Hans Effect, Glossary
- Goodhart, C. A. E. (1975). Problems of monetary management: The U.K. experience. Papers in Monetary Economics, Vol. 1. Reserve Bank of Australia.Cited in: Goodhart’s Law
- Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On calibration of modern neural networks. Proceedings of the International Conference on Machine Learning (ICML 2017).Cited in: Glossary
- Gururangan, S., Swayamdipta, S., Levy, O., Schwartz, R., Bowman, S. R., & Smith, N. A. (2018). Annotation artifacts in natural language inference data. Proceedings of NAACL-HLT 2018.Cited in: The Clever Hans Effect
- Hand, D. J. (2006). Classifier technology and the illusion of progress. Statistical Science, 21(1), 1-14.Cited in: Distribution Shift, The Baseline Rule
- Holstein, K., Wortman Vaughan, J., Daumé III, H., Dudík, M., & Wallach, H. (2019). Improving fairness in machine learning systems: What do industry practitioners need? Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI 2019).Cited in: Being Pragmatic
- Hutchinson, B., et al. (2022). Evaluation gaps in machine learning practice. Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAccT 2022).Cited in: The Functionality Fallacy, The Lab-to-Field Gap, Overview
- Jacobs, A. Z., & Wallach, H. (2021). Measurement and fairness. Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAccT 2021).Cited in: The Construct Gap
- Kapoor, S., & Narayanan, A. (2023). Leakage and the reproducibility crisis in machine-learning-based science. Patterns, 4(9), 100804.Cited in: Data Contamination
- Kapoor, S., Stroebl, B., Siegel, Z. S., Nadgir, N., & Narayanan, A. (2025). AI agents that matter. Transactions on Machine Learning Research (TMLR).Cited in: The Reproducibility Rule, Adaptive Overfitting, The Cost Frontier, Overview
- Karpinska, M., Akoury, N., & Krishna, K. (2021). The perils of using Mechanical Turk to evaluate open-ended text generation. Proceedings of EMNLP 2021.Cited in: The Gold Standard Myth
- Kiela, D., et al. (2021). Dynabench: Rethinking benchmarking in NLP. Proceedings of NAACL-HLT 2021.Cited in: Benchmark Saturation, Overview
- Koh, P. W., et al. (2021). WILDS: A benchmark of in-the-wild distribution shifts. Proceedings of the International Conference on Machine Learning (ICML 2021).Cited in: Distribution Shift
- Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. (Book)Cited in: Being Pragmatic, Glossary
- Lapuschkin, S., Wäldchen, S., Binder, A., Montavon, G., Samek, W., & Müller, K.-R. (2019). Unmasking Clever Hans predictors and assessing what machines really learn. Nature Communications, 10, 1096.Cited in: The Clever Hans Effect
- Liang, P., et al. (2023). Holistic evaluation of language models. Transactions on Machine Learning Research (TMLR).Cited in: The Benchmark Lottery, The Metric Mirage, The Cost Frontier
- Liao, Q. V., & Xiao, Z. (2023). Rethinking model evaluation as narrowing the socio-technical gap. arXiv:2306.03100. (Preprint)Cited in: The Lab-to-Field Gap, Criteria Drift
- Lin, J. (2019). The neural hype and comparisons against weak baselines. ACM SIGIR Forum, 52(2), 40-51.Cited in: The Baseline Rule
- Lipton, Z. C., & Steinhardt, J. (2019). Troubling trends in machine learning scholarship. ACM Queue, 17(1).Cited in: Campbell’s Law
- Madaio, M. A., Stark, L., Wortman Vaughan, J., & Wallach, H. (2020). Co-designing checklists to understand organizational challenges and opportunities around fairness in AI. Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI 2020).Cited in: Being Pragmatic
- Madaio, M., Egede, L., Subramonyam, H., Wortman Vaughan, J., & Wallach, H. (2022). Assessing the fairness of AI systems: AI practitioners’ processes, challenges, and needs for support. Proceedings of the ACM on Human-Computer Interaction, 6(CSCW1), Article 52.Cited in: Being Pragmatic
- Manheim, D., & Garrabrant, S. (2018). Categorizing variants of Goodhart’s Law. arXiv:1803.04585. (Preprint)Cited in: Goodhart’s Law
- McCoy, R. T., Pavlick, E., & Linzen, T. (2019). Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference. Proceedings of ACL 2019.Cited in: The Clever Hans Effect
- Melis, G., Dyer, C., & Blunsom, P. (2018). On the state of the art of evaluation in neural language models. International Conference on Learning Representations (ICLR 2018).Cited in: The Baseline Rule
- Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons’ responses and performances as scientific inquiry into score meaning. American Psychologist, 50(9), 741-749.Cited in: The Construct Gap, Overview, Glossary
- Miller, E. (2024). Adding error bars to evals: A statistical approach to language model evaluations. arXiv:2411.00640. (Preprint)Cited in: No Error Bars, No Result, Playbook
- Mitchell, M., et al. (2019). Model cards for model reporting. Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT* 2019).Cited in: The Reproducibility Rule, The Averaging Trap, Overview, Playbook, Glossary
- Mizrahi, M., et al. (2024). State of what art? A call for multi-prompt LLM evaluation. Transactions of the Association for Computational Linguistics, 12.Cited in: Prompt Sensitivity
- Moravec, H. (1988). Mind children: The future of robot and human intelligence. Harvard University Press. (Book)Cited in: The Jagged Frontier
- National Institute of Standards and Technology (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1. U.S. Department of Commerce. (Report)Cited in: The Functionality Fallacy, Overview, Being Pragmatic, Glossary
- Needham, J., Edkins, G., Pimpale, G., Bartsch, H., & Hobbhahn, M. (2025). Large language models often know when they are being evaluated. arXiv:2505.23836. (Preprint)Cited in: Evaluation Awareness, Overview, Glossary
- Northcutt, C. G., Athalye, A., & Mueller, J. (2021). Pervasive label errors in test sets destabilize machine learning benchmarks. NeurIPS 2021 Datasets and Benchmarks Track.Cited in: The Gold Standard Myth, Benchmark Saturation
- Novikova, J., Dušek, O., Cercas Curry, A., & Rieser, V. (2017). Why we need new evaluation metrics for NLG. Proceedings of EMNLP 2017.Cited in: The Metric Mirage
- Oakden-Rayner, L., Dunnmon, J., Carneiro, G., & Ré, C. (2020). Hidden stratification causes clinically meaningful failures in machine learning for medical imaging. Proceedings of the ACM Conference on Health, Inference, and Learning (CHIL 2020).Cited in: The Averaging Trap
- Ott, S., Barbosa-Silva, A., Blagec, K., Brauner, J., & Samwald, M. (2022). Mapping global dynamics of benchmark creation and saturation in artificial intelligence. Nature Communications, 13, 6793.Cited in: Benchmark Saturation, Overview
- Pan, A., Bhatia, K., & Steinhardt, J. (2022). The effects of reward misspecification: Mapping and mitigating misaligned models. International Conference on Learning Representations (ICLR 2022).Cited in: Goodhart’s Law
- Panickssery, A., Bowman, S. R., & Feng, S. (2024). LLM evaluators recognize and favor their own generations. Advances in Neural Information Processing Systems (NeurIPS 2024).Cited in: Judge Bias
- Parasuraman, R., & Riley, V. (1997). Humans and automation: Use, misuse, disuse, abuse. Human Factors, 39(2), 230-253.Cited in: The Team Is the System, Glossary
- Parnin, C., Soares, G., Pandita, R., Gulwani, S., Rich, J., & Henley, A. Z. (2025). Building your own product copilot: Challenges, opportunities, and needs. Proceedings of the IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER 2025).Cited in: Being Pragmatic
- Pineau, J., et al. (2021). Improving reproducibility in machine learning research (A report from the NeurIPS 2019 Reproducibility Program). Journal of Machine Learning Research, 22(164), 1-20.Cited in: The Reproducibility Rule
- Rabanser, S., Kapoor, S., Kirgis, P., Liu, K., Utpala, S., & Narayanan, A. (2026). Towards a science of AI agent reliability. arXiv:2602.16666. (Preprint)Cited in: Once Is Not Reliable
- Raji, I. D., Bender, E. M., Paullada, A., Denton, E., & Hanna, A. (2021). AI and the everything in the whole wide world benchmark. NeurIPS 2021 Datasets and Benchmarks Track.Cited in: The Construct Gap, Overview
- Raji, I. D., Kumar, I. E., Horowitz, A., & Selbst, A. (2022). The fallacy of AI functionality. Proceedings of the ACM Conference on Fairness, Accountability, and Transparency (FAccT 2022).Cited in: The Functionality Fallacy, Overview
- Raji, I. D., Smart, A., White, R. N., Mitchell, M., Gebru, T., et al. (2020). Closing the AI accountability gap: Defining an end-to-end framework for internal algorithmic auditing. Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT* 2020).Cited in: Being Pragmatic
- Rakova, B., Yang, J., Cramer, H., & Chowdhury, R. (2021). Where responsible AI meets reality: Practitioner perspectives on enablers for shifting organizational practices. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW1).Cited in: Being Pragmatic
- Recht, B., Roelofs, R., Schmidt, L., & Shankar, V. (2019). Do ImageNet classifiers generalize to ImageNet? Proceedings of the International Conference on Machine Learning (ICML 2019).Cited in: Adaptive Overfitting, Distribution Shift
- Reiter, E. (2018). A structured review of the validity of BLEU. Computational Linguistics, 44(3), 393-401.Cited in: The Metric Mirage
- Reuel, A., Hardy, A., Smith, C., Lamparth, M., Hardy, M., & Kochenderfer, M. J. (2024). BetterBench: Assessing AI benchmarks, uncovering issues, and establishing best practices. NeurIPS 2024 Datasets and Benchmarks Track.Cited in: No Error Bars, No Result, The Reproducibility Rule
- Ribeiro, M. T., Wu, T., Guestrin, C., & Singh, S. (2020). Beyond accuracy: Behavioral testing of NLP models with CheckList. Proceedings of ACL 2020.Cited in: The Clever Hans Effect, Overview, Glossary
- Roelofs, R., Shankar, V., Recht, B., Fridovich-Keil, S., Hardt, M., Miller, J., & Schmidt, L. (2019). A meta-analysis of overfitting in machine learning. Advances in Neural Information Processing Systems (NeurIPS 2019).Cited in: Adaptive Overfitting
- Rogers, E. M. (2003). Diffusion of innovations (5th ed.). Free Press. (Book)Cited in: Being Pragmatic
- Sainz, O., Campos, J. A., García-Ferrero, I., Etxaniz, J., Lopez de Lacalle, O., & Agirre, E. (2023). NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark. Findings of EMNLP 2023.Cited in: Data Contamination, Overview
- Salaudeen, O., et al. (2025). Measurement to meaning: A validity-centered framework for AI evaluation. arXiv:2505.10573. (Preprint)Cited in: The Construct Gap
- Schaeffer, R., Miranda, B., & Koyejo, S. (2023). Are emergent abilities of large language models a mirage? Advances in Neural Information Processing Systems (NeurIPS 2023).Cited in: The Metric Mirage
- Sclar, M., Choi, Y., Tsvetkov, Y., & Suhr, A. (2024). Quantifying language models’ sensitivity to spurious features in prompt design or: How I learned to start worrying about prompt formatting. International Conference on Learning Representations (ICLR 2024).Cited in: Prompt Sensitivity, Playbook
- Selbst, A. D., boyd, d., Friedler, S. A., Venkatasubramanian, S., & Vertesi, J. (2019). Fairness and abstraction in sociotechnical systems. Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT* 2019).Cited in: The Lab-to-Field Gap
- Shankar, S., Garcia, R., Hellerstein, J. M., & Parameswaran, A. G. (2024). “We have no idea how models will behave in production until production”: How engineers operationalize machine learning. Proceedings of the ACM on Human-Computer Interaction, 8(CSCW1), Article 206.Cited in: Being Pragmatic
- Shankar, S., Zamfirescu-Pereira, J. D., Hartmann, B., Parameswaran, A. G., & Arawjo, I. (2024). Who validates the validators? Aligning LLM-assisted evaluation of LLM outputs with human preferences. Proceedings of the ACM Symposium on User Interface Software and Technology (UIST 2024).Cited in: Criteria Drift, Judge Bias, Being Pragmatic, Playbook
- Shevlane, T., et al. (2023). Model evaluation for extreme risks. arXiv:2305.15324. (Preprint)Cited in: Presence, Not Absence, Overview
- Singh, S., Nan, Y., Wang, A., D’Souza, D., Kapoor, S., et al. (2025). The leaderboard illusion. Advances in Neural Information Processing Systems (NeurIPS 2025).Cited in: Goodhart’s Law, Campbell’s Law
- Strathern, M. (1997). ‘Improving ratings’: Audit in the British university system. European Review, 5(3), 305-321.Cited in: Goodhart’s Law
- Vaccaro, M., Almaatouq, A., & Malone, T. (2024). When combinations of humans and AI are useful: A systematic review and meta-analysis. Nature Human Behaviour, 8, 2293-2303.Cited in: The Team Is the System, The Jagged Frontier, Playbook
- van der Lee, C., Gatt, A., van Miltenburg, E., & Krahmer, E. (2021). Human evaluation of automatically generated text: Current trends and best practice guidelines. Computer Speech & Language, 67, 101151.Cited in: The Gold Standard Myth, Overview
- van der Maden, W., Sadek, M., Xiao, Z., Mottelson, A., Liao, Q. V., & Zhu, J. (2026). Results-actionability gap: Understanding how practitioners evaluate LLM products in the wild. Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI 2026).Cited in: Being Pragmatic, Glossary
- van der Weij, T., Hofstätter, F., Jaffe, O., Brown, S. F., & Ward, F. R. (2025). AI sandbagging: Language models can strategically underperform on evaluations. International Conference on Learning Representations (ICLR 2025).Cited in: Presence, Not Absence, Evaluation Awareness, Glossary
- Wallach, H., Desai, M., Cooper, A. F., Wang, A., Atalla, C., et al. (2025). Position: Evaluating generative AI systems is a social science measurement challenge. Proceedings of the International Conference on Machine Learning (ICML 2025), PMLR 267.Cited in: The Construct Gap, Overview, Being Pragmatic
- Wang, P., et al. (2024). Large language models are not fair evaluators. Proceedings of ACL 2024.Cited in: Judge Bias
- Weidinger, L., Rauh, M., Marchal, N., Manzini, A., et al. (2023). Sociotechnical safety evaluation of generative AI systems. arXiv:2310.11986. (Preprint)Cited in: The Lab-to-Field Gap, Overview
- Wong, A., Otles, E., Donnelly, J. P., et al. (2021). External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients. JAMA Internal Medicine, 181(8), 1065-1070.Cited in: The Functionality Fallacy, Distribution Shift
- Wu, T., Ribeiro, M. T., Heer, J., & Weld, D. (2019). Errudite: Scalable, reproducible, and testable error analysis. Proceedings of ACL 2019.Cited in: Being Pragmatic, Glossary
- Yao, S., Shinn, N., Razavi, P., & Narasimhan, K. (2025). τ-bench: A benchmark for tool-agent-user interaction in real-world domains. International Conference on Learning Representations (ICLR 2025).Cited in: Once Is Not Reliable, Playbook, Glossary
- Zhang, H., Da, J., Lee, D., Robinson, V., Wu, C., et al. (2024). A careful examination of large language model performance on grade school arithmetic. NeurIPS 2024 Datasets and Benchmarks Track.Cited in: Data Contamination
- Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., et al. (2023). Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. NeurIPS 2023 Datasets and Benchmarks Track.Cited in: Judge Bias, Overview, Playbook, Glossary