Guide
Being Pragmatic
The patterns on this site describe the ways evaluation goes wrong. This page is about doing it anyway, inside a team with deadlines, limited budget, and a model that changed last Tuesday.
Perfect evaluation doesn’t exist. What you’re after is evidence that’s good enough for the decision in front of you, gathered in a way the team will keep doing next quarter. Research on how practitioners actually work backs this up. Teams start with informal checks, struggle to turn results into changes, and feel constant pressure to ship (van der Maden et al., 2026; Madaio et al., 2022). A pragmatic approach works with those realities instead of pretending they aren’t there.
Ten principles
1. Match the rigor to the stakes
Not every AI feature needs a full evaluation program. A tool that drafts internal meeting notes and a tool that flags patients for sepsis don’t deserve the same process. The NIST AI Risk Management Framework (2023) says policies and resources should be prioritized by risk level and potential impact, and notes that trying to eliminate all negative risk can be counterproductive. The European Union’s AI Act (2024) takes the same risk-based approach, with its heaviest requirements reserved for high-risk uses.
| Tier | Looks like | Minimum evidence before launch | Patterns to lean on |
|---|---|---|---|
| Low | Internal, easy to undo, a person reviews every output | A saved set of real examples, a team review of outputs, a named owner | The Construct Gap, Criteria Drift |
| Medium | Customer-facing, reversible, moderate cost of errors | A locked test set, a rubric, a baseline, multiple runs, one or two slices, an eval card, a staged rollout | Prompt Sensitivity, Once Is Not Reliable, The Baseline Rule, The Averaging Trap |
| High | Affects health, money, safety, rights, or access; hard to reverse | Everything above, plus validated graders, confidence intervals, subgroup analysis, red teaming, a field pilot, independent review, and post-launch monitoring | All of them, especially The Functionality Fallacy, Presence, Not Absence, and The Lab-to-Field Gap |
Engineers building LLM features have asked this question directly: where’s the line between enough evidence and overspending chasing perfection (Parnin et al., 2025)? Tiers give the team an answer they agreed on in advance, instead of relitigating it for every launch.
2. Start with vibe checks, then write them down
Every team starts by trying the thing and seeing how it feels. That’s fine. In interviews with 19 practitioners building LLM products, van der Maden et al. (2026) found informal “vibe checks” described as “irreplaceable, they are the first line of evaluation.” The problem isn’t vibe checks. It’s stopping there.
The move is to formalize gradually. Save the examples you tried. Write down what you were looking for. When the same note shows up three times, it’s a criterion. Madaio et al. (2020) found that checklists can give teams infrastructure for formalizing ad hoc processes, as long as they’re grounded in practitioners’ real needs. Otherwise they get misused.
3. Tie every eval to a decision
Van der Maden et al. (2026) named the most common failure the results-actionability gap: teams collect evaluation data but can’t turn it into concrete improvements. It affected 17 of their 19 participants. One put it plainly: “When you have an evaluation scale and you end up with a 4.6 out of ten, if you don’t know what caused that, then it’s very difficult to iterate on making it better.”
Two habits help:
- Write the decision rule first. “If accuracy on refund requests is below 90%, we don’t expand the rollout.” “If the new prompt doesn’t beat the old one by more than the error bars, we keep the old one.”
- Change one thing at a time. Prompt, model, retrieval settings, and temperature all interact. If you change three at once, a score change tells you almost nothing about what to fix.
4. Fit into rituals the team already has
New processes that feel like extra work die quietly. Rogers (2003) identified what speeds up adoption of a new practice: it’s clearly better than what people do now, compatible with how they already work, simple to understand, easy to try on a small scale, and visible to others. Evaluation is no different.
- Add 15 minutes of output review to an existing sprint review or design critique instead of creating a new meeting.
- Put the eval card in the launch review template that already exists.
- Run a small canary eval in the same pipeline that runs other tests.
- Share one interesting failure in the team channel each week. It makes the work visible without a status report.
5. Look at outputs together
The fastest way to build shared judgment is to read real outputs as a group: product, design, engineering, and someone with domain expertise. Disagreements about whether an output is “good” surface the definitions you haven’t written yet (Wallach et al., 2025; Shankar et al., 2024).
Keep the session honest. Wu et al. (2019) found that the common practice of eyeballing a small sample of errors can lead to biased, incomplete conclusions. They recommend defining error groups precisely, looking at enough examples (including successes, not just failures), and testing your hypothesis about what causes an error instead of assuming it.
Bring in domain experts early. Madaio et al. (2022) found that lack of engagement with stakeholders and domain experts was one of the main organizational barriers to meaningful evaluation.
6. Give it an owner, not a police force
When evaluation is everyone’s job, it’s no one’s. Rakova et al. (2021) interviewed practitioners across 19 organizations and found role uncertainty was a recurring barrier. As one put it, “Whose job is this?” Work fell through the cracks, depended on individual volunteers, and was hard to credit in performance reviews.
- Name an owner for each AI feature’s evaluation. On a small team, that can be a hat someone wears, not a new role.
- Put evaluation work in role expectations and recognize it in reviews.
- The owner’s job is to make evaluation easy for the team, not to approve or reject other people’s work.
7. Make bad news safe to share
An eval that finds a problem is doing its job. If finding one feels like getting someone in trouble, people stop looking. Edmondson (1999) found that psychological safety, “a shared belief that the team is safe for interpersonal risk taking,” was linked to learning behavior in work teams.
- Celebrate failures found before launch. That’s the cheapest time to find them.
- Run blameless reviews when something slips through. Google’s site reliability practice treats postmortems as a way to learn from incidents, not to assign blame (Beyer et al., 2016).
- Leaders go first: share a result that didn’t go the way you hoped.
8. Spend the budget where it counts
Evaluation costs money and time, and the costs add up. Engineers building product copilots described tests that cost a cent or two each, which became real money at scale, and one was asked to stop running benchmarks because of the cost (Parnin et al., 2025).
Layer it:
- Every change: a small, fast canary set with automated checks.
- Every milestone: the full locked test set with repeated runs and slices.
- High-stakes slices: human review, every time.
- Before a big launch: red teaming and a field pilot, scaled to the tier.
9. Assume production will surprise you
The title of one interview study says it: “We have no idea how models will behave in production until production” (Shankar et al., 2024). The engineers in that study evaluated throughout a multi-stage deployment and kept monitoring after launch, balancing velocity against visibility and versioning.
That’s not an excuse to skip pre-launch testing. It’s a reason to plan for what happens after.
- Roll out in stages. Internal users, then a small percentage of customers, then more. Controlled online experiments let you measure real impact while limiting exposure (Kohavi et al., 2020).
- Set an error budget. Borrowed from site reliability engineering (Beyer et al., 2016): agree ahead of time on an acceptable failure rate, and on what the team does when it’s exceeded.
- Don’t wait for complaints. In a survey of industry practitioners, about half said their teams had found serious fairness issues only after deploying a system. One engineer described the default as putting the model out there, and “then you’ll know if there’s fairness issues if someone raises hell online” (Holstein et al., 2019). Internal audits across the development lifecycle are the proactive alternative (Raji et al., 2020).
10. Use the patterns as questions, not weapons
Nobody wants the coworker who quotes Goodhart’s Law in every meeting. The patterns work best as questions, asked at the right moment, about the two or three risks that matter for the decision at hand.
- Instead of “That’s the Construct Gap,” try “What would a user need to be able to do for this score to mean what we want?”
- Instead of “No error bars, no result,” try “How much would this number move if we ran it again?”
- Instead of “That’s contamination,” try “Could the model have seen these questions before?”
And don’t let the perfect eval block a good one. Breck et al. (2017) framed production readiness as a rubric teams score themselves against and improve over time. Treat evaluation maturity the same way.
A maturity path
Most teams move through these stages. Van der Maden et al. (2026) describe it as a formalization journey. Skipping ahead rarely sticks, so aim for the next stage, not the last one.
| Stage | What it looks like | Next small step |
|---|---|---|
| 0. Vibes | People try it and share impressions | Save the examples you tried and what you noticed |
| 1. Saved examples | A shared doc or spreadsheet of real inputs and outputs | Write a one-sentence construct and a first rubric |
| 2. Repeatable eval | A locked test set, a rubric, a baseline, a person responsible | Add multiple prompts and runs, and one slice |
| 3. Built into the workflow | Automated canary checks, human spot checks, intervals, slices, eval cards in launch reviews | Add staged rollouts and production monitoring |
| 4. Continuous | Monitoring, error budgets, periodic audits, rubrics that evolve with version history | Share what you learned with other teams |
A starter plan for the first 30 days
Week 1. Look. Pick one AI feature. Pull 30 to 50 real examples. Run a one-hour output review with product, design, engineering, and a domain expert. Note what’s good, what’s bad, and what people disagree about.
Week 2. Define. Write the construct in one sentence. Draft a rubric from the week 1 notes. Pick a baseline. Write one decision rule.
Week 3. Measure. Build a small locked test set. Run it with three prompt variants and three runs each. Compute simple confidence intervals. Break results down by one slice that matters.
Week 4. Embed. Add an eval card to the launch review. Name an owner. Set a rerun trigger, like every model or prompt change. Share one finding with the wider team.
Who does what
Small teams will combine these. What matters is that each job has a name next to it.
| Role | Owns |
|---|---|
| Product | The decision, the risk tier, and the thresholds |
| Design and UX research | Construct definitions, rubrics, user and field studies, human-AI team evaluation |
| Engineering and ML | The eval harness, versioning, automation, monitoring |
| Domain experts | Labels, edge cases, rubric review |
| Leadership | Time and budget, recognition, and making evaluation part of launch criteria |
Common pushback, and how to answer it
| You hear | Try |
|---|---|
| “We don’t have time.” | Start with 30 examples and one hour. Finding the problem after launch costs more. |
| “The model changes every month anyway.” | That’s why you want a small canary set you can rerun in minutes. |
| “Our users will tell us if something’s wrong.” | Some will, often late and in public. User feedback is one input, not the whole picture. |
| “The benchmark says this model is the best.” | Best at that benchmark. Let’s check it on our data. |
| “Evaluation will slow us down.” | Match the effort to the tier, and ship in stages so you learn while you ship. |
| “Quality is subjective. You can’t measure it.” | Then let’s define what we mean, write it down, and see whether two people agree. |
| “The LLM judge is good enough.” | It might be. Let’s check it against 50 human labels first. |
References
- Beyer, B., Jones, C., Petoff, J., & Murphy, N. R. (Eds.) (2016). Site reliability engineering: How Google runs production systems. O’Reilly Media. (Book)
- Breck, E., Cai, S., Nielsen, E., Salib, M., & Sculley, D. (2017). The ML test score: A rubric for ML production readiness and technical debt reduction. Proceedings of the IEEE International Conference on Big Data.
- Edmondson, A. (1999). Psychological safety and learning behavior in work teams. Administrative Science Quarterly, 44(2), 350-383.
- European Union (2024). Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union. (Report)
- Holstein, K., Wortman Vaughan, J., Daumé III, H., Dudík, M., & Wallach, H. (2019). Improving fairness in machine learning systems: What do industry practitioners need? Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI 2019).
- Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. (Book)
- Madaio, M. A., Stark, L., Wortman Vaughan, J., & Wallach, H. (2020). Co-designing checklists to understand organizational challenges and opportunities around fairness in AI. Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI 2020).
- Madaio, M., Egede, L., Subramonyam, H., Wortman Vaughan, J., & Wallach, H. (2022). Assessing the fairness of AI systems: AI practitioners’ processes, challenges, and needs for support. Proceedings of the ACM on Human-Computer Interaction, 6(CSCW1), Article 52.
- National Institute of Standards and Technology (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1. U.S. Department of Commerce. (Report)
- Parnin, C., Soares, G., Pandita, R., Gulwani, S., Rich, J., & Henley, A. Z. (2025). Building your own product copilot: Challenges, opportunities, and needs. Proceedings of the IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER 2025).
- Raji, I. D., Smart, A., White, R. N., Mitchell, M., Gebru, T., et al. (2020). Closing the AI accountability gap: Defining an end-to-end framework for internal algorithmic auditing. Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT* 2020).
- Rakova, B., Yang, J., Cramer, H., & Chowdhury, R. (2021). Where responsible AI meets reality: Practitioner perspectives on enablers for shifting organizational practices. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW1).
- Rogers, E. M. (2003). Diffusion of innovations (5th ed.). Free Press. (Book)
- Shankar, S., Garcia, R., Hellerstein, J. M., & Parameswaran, A. G. (2024). “We have no idea how models will behave in production until production”: How engineers operationalize machine learning. Proceedings of the ACM on Human-Computer Interaction, 8(CSCW1), Article 206.
- Shankar, S., Zamfirescu-Pereira, J. D., Hartmann, B., Parameswaran, A. G., & Arawjo, I. (2024). Who validates the validators? Aligning LLM-assisted evaluation of LLM outputs with human preferences. Proceedings of the ACM Symposium on User Interface Software and Technology (UIST 2024).
- van der Maden, W., Sadek, M., Xiao, Z., Mottelson, A., Liao, Q. V., & Zhu, J. (2026). Results-actionability gap: Understanding how practitioners evaluate LLM products in the wild. Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI 2026).
- Wallach, H., Desai, M., Cooper, A. F., Wang, A., Atalla, C., et al. (2025). Position: Evaluating generative AI systems is a social science measurement challenge. Proceedings of the International Conference on Machine Learning (ICML 2025), PMLR 267.
- Wu, T., Ribeiro, M. T., Heer, J., & Weld, D. (2019). Errudite: Scalable, reproducible, and testable error analysis. Proceedings of ACL 2019.