Guide

Being Pragmatic

The patterns on this site describe the ways evaluation goes wrong. This page is about doing it anyway, inside a team with deadlines, limited budget, and a model that changed last Tuesday.

Perfect evaluation doesn’t exist. What you’re after is evidence that’s good enough for the decision in front of you, gathered in a way the team will keep doing next quarter. Research on how practitioners actually work backs this up. Teams start with informal checks, struggle to turn results into changes, and feel constant pressure to ship (van der Maden et al., 2026; Madaio et al., 2022). A pragmatic approach works with those realities instead of pretending they aren’t there.

Ten principles

1. Match the rigor to the stakes

Not every AI feature needs a full evaluation program. A tool that drafts internal meeting notes and a tool that flags patients for sepsis don’t deserve the same process. The NIST AI Risk Management Framework (2023) says policies and resources should be prioritized by risk level and potential impact, and notes that trying to eliminate all negative risk can be counterproductive. The European Union’s AI Act (2024) takes the same risk-based approach, with its heaviest requirements reserved for high-risk uses.

TierLooks likeMinimum evidence before launchPatterns to lean on
LowInternal, easy to undo, a person reviews every outputA saved set of real examples, a team review of outputs, a named ownerThe Construct Gap, Criteria Drift
MediumCustomer-facing, reversible, moderate cost of errorsA locked test set, a rubric, a baseline, multiple runs, one or two slices, an eval card, a staged rolloutPrompt Sensitivity, Once Is Not Reliable, The Baseline Rule, The Averaging Trap
HighAffects health, money, safety, rights, or access; hard to reverseEverything above, plus validated graders, confidence intervals, subgroup analysis, red teaming, a field pilot, independent review, and post-launch monitoringAll of them, especially The Functionality Fallacy, Presence, Not Absence, and The Lab-to-Field Gap

Engineers building LLM features have asked this question directly: where’s the line between enough evidence and overspending chasing perfection (Parnin et al., 2025)? Tiers give the team an answer they agreed on in advance, instead of relitigating it for every launch.

2. Start with vibe checks, then write them down

Every team starts by trying the thing and seeing how it feels. That’s fine. In interviews with 19 practitioners building LLM products, van der Maden et al. (2026) found informal “vibe checks” described as “irreplaceable, they are the first line of evaluation.” The problem isn’t vibe checks. It’s stopping there.

The move is to formalize gradually. Save the examples you tried. Write down what you were looking for. When the same note shows up three times, it’s a criterion. Madaio et al. (2020) found that checklists can give teams infrastructure for formalizing ad hoc processes, as long as they’re grounded in practitioners’ real needs. Otherwise they get misused.

3. Tie every eval to a decision

Van der Maden et al. (2026) named the most common failure the results-actionability gap: teams collect evaluation data but can’t turn it into concrete improvements. It affected 17 of their 19 participants. One put it plainly: “When you have an evaluation scale and you end up with a 4.6 out of ten, if you don’t know what caused that, then it’s very difficult to iterate on making it better.”

Two habits help:

4. Fit into rituals the team already has

New processes that feel like extra work die quietly. Rogers (2003) identified what speeds up adoption of a new practice: it’s clearly better than what people do now, compatible with how they already work, simple to understand, easy to try on a small scale, and visible to others. Evaluation is no different.

5. Look at outputs together

The fastest way to build shared judgment is to read real outputs as a group: product, design, engineering, and someone with domain expertise. Disagreements about whether an output is “good” surface the definitions you haven’t written yet (Wallach et al., 2025; Shankar et al., 2024).

Keep the session honest. Wu et al. (2019) found that the common practice of eyeballing a small sample of errors can lead to biased, incomplete conclusions. They recommend defining error groups precisely, looking at enough examples (including successes, not just failures), and testing your hypothesis about what causes an error instead of assuming it.

Bring in domain experts early. Madaio et al. (2022) found that lack of engagement with stakeholders and domain experts was one of the main organizational barriers to meaningful evaluation.

6. Give it an owner, not a police force

When evaluation is everyone’s job, it’s no one’s. Rakova et al. (2021) interviewed practitioners across 19 organizations and found role uncertainty was a recurring barrier. As one put it, “Whose job is this?” Work fell through the cracks, depended on individual volunteers, and was hard to credit in performance reviews.

7. Make bad news safe to share

An eval that finds a problem is doing its job. If finding one feels like getting someone in trouble, people stop looking. Edmondson (1999) found that psychological safety, “a shared belief that the team is safe for interpersonal risk taking,” was linked to learning behavior in work teams.

8. Spend the budget where it counts

Evaluation costs money and time, and the costs add up. Engineers building product copilots described tests that cost a cent or two each, which became real money at scale, and one was asked to stop running benchmarks because of the cost (Parnin et al., 2025).

Layer it:

9. Assume production will surprise you

The title of one interview study says it: “We have no idea how models will behave in production until production” (Shankar et al., 2024). The engineers in that study evaluated throughout a multi-stage deployment and kept monitoring after launch, balancing velocity against visibility and versioning.

That’s not an excuse to skip pre-launch testing. It’s a reason to plan for what happens after.

10. Use the patterns as questions, not weapons

Nobody wants the coworker who quotes Goodhart’s Law in every meeting. The patterns work best as questions, asked at the right moment, about the two or three risks that matter for the decision at hand.

And don’t let the perfect eval block a good one. Breck et al. (2017) framed production readiness as a rubric teams score themselves against and improve over time. Treat evaluation maturity the same way.

A maturity path

Most teams move through these stages. Van der Maden et al. (2026) describe it as a formalization journey. Skipping ahead rarely sticks, so aim for the next stage, not the last one.

StageWhat it looks likeNext small step
0. VibesPeople try it and share impressionsSave the examples you tried and what you noticed
1. Saved examplesA shared doc or spreadsheet of real inputs and outputsWrite a one-sentence construct and a first rubric
2. Repeatable evalA locked test set, a rubric, a baseline, a person responsibleAdd multiple prompts and runs, and one slice
3. Built into the workflowAutomated canary checks, human spot checks, intervals, slices, eval cards in launch reviewsAdd staged rollouts and production monitoring
4. ContinuousMonitoring, error budgets, periodic audits, rubrics that evolve with version historyShare what you learned with other teams

A starter plan for the first 30 days

Week 1. Look. Pick one AI feature. Pull 30 to 50 real examples. Run a one-hour output review with product, design, engineering, and a domain expert. Note what’s good, what’s bad, and what people disagree about.

Week 2. Define. Write the construct in one sentence. Draft a rubric from the week 1 notes. Pick a baseline. Write one decision rule.

Week 3. Measure. Build a small locked test set. Run it with three prompt variants and three runs each. Compute simple confidence intervals. Break results down by one slice that matters.

Week 4. Embed. Add an eval card to the launch review. Name an owner. Set a rerun trigger, like every model or prompt change. Share one finding with the wider team.

Who does what

Small teams will combine these. What matters is that each job has a name next to it.

RoleOwns
ProductThe decision, the risk tier, and the thresholds
Design and UX researchConstruct definitions, rubrics, user and field studies, human-AI team evaluation
Engineering and MLThe eval harness, versioning, automation, monitoring
Domain expertsLabels, edge cases, rubric review
LeadershipTime and budget, recognition, and making evaluation part of launch criteria

Common pushback, and how to answer it

You hearTry
“We don’t have time.”Start with 30 examples and one hour. Finding the problem after launch costs more.
“The model changes every month anyway.”That’s why you want a small canary set you can rerun in minutes.
“Our users will tell us if something’s wrong.”Some will, often late and in public. User feedback is one input, not the whole picture.
“The benchmark says this model is the best.”Best at that benchmark. Let’s check it on our data.
“Evaluation will slow us down.”Match the effort to the tier, and ship in stages so you learn while you ship.
“Quality is subjective. You can’t measure it.”Then let’s define what we mean, write it down, and see whether two people agree.
“The LLM judge is good enough.”It might be. Let’s check it against 50 human labels first.

References