About
About this guide
An independent reference for people who build, buy, or design with AI. It collects the reliable ways AI evaluation goes wrong, with the research behind each one.
What a “pattern” means here
A pattern here is a reliable way evaluation goes wrong, with research behind it. It is not a law of physics. Some have established names, like Goodhart’s Law. Others are named on this site to make a well-documented idea easier to remember, and each page says which.
Editorial policy
- Every pattern links to its sources. Preprints, reports, working papers, and books are labeled as such. Everything else is peer-reviewed or a classic in its field.
- Each pattern shows when its evidence was last checked and carries a version number and revision history.
- The Changelog records what changed and when. Corrections are made in the open.
- Role tags, plain-terms summaries, checklist questions, and claim-checker rules are editorial work built from the patterns themselves. They are starting points, not authority.
Status of this guide
This is an independent, self-published reference built from published research. It has not been through formal peer review. Every pattern links to its sources, so you can check the evidence for yourself.
Established terms and names from this site
Not every pattern name is standard. The idea behind each pattern is established in the research, but for some the name is this site’s own.
- Established terms (13): Goodhart’s Law; Campbell’s Law; The Functionality Fallacy; Benchmark Saturation; Data Contamination; Adaptive Overfitting; The Clever Hans Effect; Prompt Sensitivity; Criteria Drift; The Benchmark Lottery; Evaluation Awareness; Distribution Shift; The Jagged Frontier.
- Named on this site (13): The Construct Gap; The Metric Mirage; The Gold Standard Myth; Judge Bias; Once Is Not Reliable; The Cost Frontier; No Error Bars, No Result; The Averaging Trap; The Baseline Rule; The Reproducibility Rule; Presence, Not Absence; The Lab-to-Field Gap; The Team Is the System. Please do not cite these as established terms. Cite the sources linked on each page instead.
Review
See the methodology for how sources were chosen and checked. No outside reviewer has signed off on any pattern yet. If you work in evaluation, statistics, or a related field and see something wrong, or a better source, please say so on the contact page. Corrections are recorded in the Changelog.
Who maintains it
Hi, I’m Joseph Alfonso, a UX design lead. I made this guide because I wanted to understand how AI really gets judged, and how those judgments go wrong. Writing it down was how I learned.
I care about this because AI can feel like something that happens to us. I don’t think it has to. The more we understand how it works and why it fails, the more we can use it on our own terms and ask better questions of the people selling it, building it, or telling us to trust it.
I’m still learning, and I’ve tried to be honest about what I know and don’t. If you find a mistake, or a better source, please tell me through the contact page. I’d be grateful, and I’ll fix it. I hope this helps you the way making it helped me.
How to use it
Start with the guide, browse the patterns, describe your circumstances to find your patterns, or paste a claim into the claim checker to see which patterns apply. Collect any patterns with the Add buttons to get a quick brief, or turn them into a printable sheet with the checklist builder. For ready-made question lists, see buying an AI model or vendor, a launch review, and building an evaluation.