Tool
Find your patterns
Describe where you are with an AI system: what you are building or buying, how you test it, who it affects. You get the patterns worth checking first, why each applies, and a question to ask.
What we picked up Detected from your description, edit if wrong
Runs in your browser; nothing you type is sent anywhere. Matching is rule-based. How matching works
Results
Results appear here. Describe your situation, or choose what applies, then choose Find my patterns.
How matching works
Your description is scanned for situations a careful reader would recognize, such as buying a model, using an AI grader, or launching to real users. You can add or remove any of them. Each situation maps to the patterns that matter most for it, using the manuscript’s Playbook and Being Pragmatic. Patterns are ranked by how many of your situations point to them and how strongly: the top five are where to start, the next six are likely to matter, and the rest are worth a look. This is a reading aid, not an assessment of your system.
Show the rules
Choosing a model or vendor
- The Functionality Fallacy strong match You are deciding whether to rely on a system you did not build. Before comparing options, ask for evidence that any of them works on data like yours.
- The Construct Gap strong match Vendors usually show scores on tests with impressive names. Check that those tests resemble the work you need done.
- The Baseline Rule likely A vendor’s win only counts against a fair comparison. Ask what it beat, and whether your current process was in the comparison.
- The Benchmark Lottery likely Which benchmarks appear in a vendor’s table can decide who looks best. Ask which common ones are missing.
- Data Contamination likely Public benchmark scores may reflect memorized test items. Ask how overlap with training data was ruled out.
- The Cost Frontier likely Two options can score the same at very different cost. Ask for cost and latency next to every accuracy number.
- The Reproducibility Rule worth checking A vendor’s result you cannot rerun is an anecdote. Check that the setup is recorded well enough for you to repeat it.
Building our own evaluation
- Adaptive Overfitting strong match Your own test set stops being a fair test once you iterate against it. Lock a portion away before you start tuning.
- The Gold Standard Myth strong match The answer key you build is a measurement with errors. Check how often independent raters agree before you trust it.
- No Error Bars, No Result strong match A small eval set gives a wide range of plausible scores. Decide how many items and runs you need before comparing anything.
- Criteria Drift likely You will learn what good means by reading real outputs, so your rubric will change. Version it and keep a fixed reference set.
- Prompt Sensitivity likely Your results will depend on the prompt you choose. Test rephrasings before treating a score as a property of the model.
- The Metric Mirage likely Pick and justify your metrics before you see results, so the scoring method does not decide the conclusion.
- The Reproducibility Rule likely Record the model version, prompts, settings and data from the start, so the eval can be rerun later.
About to launch or pilot
- The Lab-to-Field Gap strong match Testing before launch happens in conditions that are cleaner than real use. Plan a pilot and measure what changes for real users.
- Distribution Shift strong match Your test data comes from a past population. Check how it compares with the inputs you will actually receive.
- The Averaging Trap likely A good overall score can hide a group the system fails. Decide which user groups matter and report each.
- Once Is Not Reliable likely Users meet the all-runs pass rate, not your best run. Measure how often it works every time.
- Evaluation Awareness likely Models can behave differently when they can tell they are being tested. Build launch tests from messy, real usage.
- Presence, Not Absence likely A clean pre-launch test shows you did not find a problem, not that none exists. Ask how hard anyone tried.
- The Jagged Frontier worth checking Capability is uneven. Test tasks that seem easy to people, not only the hard ones, and tell users where the known gaps are.
Already in production
- Distribution Shift strong match Real inputs change over time. Compare current production inputs with the data you tested on.
- The Lab-to-Field Gap strong match What users report may not show up in your earlier tests. Find which lab results did not hold in real use.
- Goodhart’s Law likely Once a metric guides decisions it stops being a good measure. Decide who judges when your live metric has stopped being useful.
- Evaluation Awareness likely Compare behavior in your tests with behavior in monitored real use, because the two can differ.
- The Averaging Trap likely An average over all users can hide the group that is complaining. Look at the worst-performing slice.
Using an AI to grade outputs
- Judge Bias strong match AI graders have favorites, such as longer answers or their own model family. Check the grader against human judgments on a sample.
- Criteria Drift likely The grading prompt is your rubric, and it will change as you read outputs. Version it and re-grade a fixed reference set.
- The Gold Standard Myth likely Human ratings used to check the grader also contain errors and disagreement. Measure how often raters agree.
- Prompt Sensitivity worth checking The grader is a prompt too. Small changes in how it is asked to score can move results.
Scores are noisy or inconsistent
- Once Is Not Reliable strong match If results change between runs, one run is not a result. Report the pass rate across repeated runs.
- No Error Bars, No Result strong match Differences smaller than the noise are not differences. Put an interval on each score and compare on the same items.
- Prompt Sensitivity likely Small changes in wording or formatting can swing a score. Measure how much it moves across reasonable rephrasings.
- The Reproducibility Rule likely If you cannot reproduce a result, record the model version, settings and data so you can find out why.
Affects people's decisions or safety
- The Functionality Fallacy strong match When a system affects people, whether it works comes before whether it is fair. Ask what it does when it is wrong, and how often.
- The Averaging Trap strong match A strong average can hide a group the system fails badly. Report results for the groups that matter.
- The Team Is the System likely If a person reviews the output, test the person and the AI together, not the AI alone.
- Presence, Not Absence likely An evaluation cannot prove a harm is absent. Ask who tried to find it and whether they had domain expertise.
- The Gold Standard Myth worth checking In high-stakes areas the labels are often contested. Ask who audited the items where strong models disagree with the answer key.
People work alongside the AI
- The Team Is the System strong match The person and the AI together produce the outcome. Compare the pair with the human alone and the AI alone.
- The Jagged Frontier likely People cannot see where the AI’s edges are. Tell users the known failure pockets, and check which tasks they over-trust it on.
- The Lab-to-Field Gap worth checking Real workflows add time pressure and interruptions. Measure what changes when people use it for real.
Relying on public benchmarks or leaderboards
- The Construct Gap strong match A benchmark measures what it measures, not what its name says. Check how it resembles your real tasks.
- Benchmark Saturation strong match Tests stop separating systems once everyone scores near the top. Check how close the leaders are to the ceiling.
- Data Contamination strong match Public test items can end up in training data. Ask how overlap was checked and whether newer data confirms the result.
- The Benchmark Lottery likely Which benchmarks are shown can decide who wins. Ask which common ones are missing from the table.
- Goodhart’s Law likely Systems get tuned toward well-known benchmarks. Check whether the reported metric was also the one optimized.
Many users, languages or groups
- The Averaging Trap strong match An overall score can hide the language or group that fails. Name the groups that matter and report each with its own error bars.
- Distribution Shift likely Results on one population do not automatically carry over to another. Check how your test data compares with each group you serve.
- The Jagged Frontier worth checking Capability is uneven across tasks and groups. Document the known failure pockets and tell users.
The team is chasing a number
- Goodhart’s Law strong match When a measure becomes a target it stops being a good measure. Check whether the metric you report is the one you optimized.
- Campbell’s Law strong match The more a number drives decisions, the more people bend their work around it. Back important decisions with independent measures.
- Adaptive Overfitting likely Repeated tuning against the same test set makes it worth less each time. Keep a locked set you rarely consult.
- The Metric Mirage likely The scoring method is part of the result. Check that the gain holds under a second metric.
Agents, multi-step workflows or cost
- The Cost Frontier strong match You can buy accuracy with more compute. Report cost and latency next to every accuracy number and compare at matched cost.
- Once Is Not Reliable likely Multi-step systems compound small failure rates. Measure the pass rate across repeated runs of the whole task.
- The Baseline Rule likely A complex system should beat a simple one tuned just as hard. Give simple baselines the same tuning budget.
Safety, red-teaming or 'it never does X'
- Presence, Not Absence strong match A test can show a risk exists but not that none does. Ask how hard anyone tried and with which tools.
- Evaluation Awareness likely A model that can tell it is being tested may behave differently in real use. Compare test behavior with monitored real behavior.
- Prompt Sensitivity worth checking Whether a jailbreak works can depend on small wording changes. A single attempt is not coverage.
Labels, raters or ground truth
- The Gold Standard Myth strong match Labels and human ratings are measurements with error. Measure how often independent raters agree before trusting the answer key.
- The Clever Hans Effect likely A model can match the labels for the wrong reason. Run partial-input baselines to check it uses the part of the input that should matter.
- Criteria Drift likely If the rating guidelines changed, re-grade a fixed reference set so trends stay comparable.
Suspiciously good results
- The Clever Hans Effect strong match A model can get the right answer for the wrong reason. Check whether accuracy holds when the part that should matter is removed.
- Data Contamination strong match If the test was in the training data, the score measures memory. Ask what overlap analysis was run and confirm on newer data.
- Adaptive Overfitting likely Repeated tuning against the same test makes it flatter you. Check for a locked test set that was rarely used.