Guide
Methodology
How the patterns were chosen, what counts as a source, how the evidence was checked, and what this guide does not claim to be.
How patterns are chosen
The 26 patterns are an editorial selection of recurring ways AI evaluation goes wrong, each with published research behind it. They are not the result of a systematic review, and another reasonable list would differ. Each pattern page says whether its name is an established term or this site’s own name for an established idea. The full list is on the About page.
What counts as a source
The guide cites 114 sources across all patterns. Of these, 92 are peer-reviewed or classic works, 15 are preprints, 5 are reports, 1 is a working paper, and 1 is a book. Anything that is not peer reviewed or a classic in its field is labeled on the pattern page. The Evidence section reports what the authors found, and where results are mixed or limited the page says so, for example in Adaptive Overfitting and Data Contamination. Every pattern also has a “Where this doesn’t apply” section.
How the evidence was checked
Each pattern shows when its evidence was last reviewed. In October 2026, more than 40 specific figures and findings on the pattern pages were compared with the source’s abstract or full text, covering at least one claim in 25 of the 26 patterns. No discrepancies were found. Every source link was also tested, and every DOI resolves.
The limits of that check: it covered numbers and the main finding of each source, not whether every interpretation or piece of advice goes beyond the source. A few details were confirmed through secondary descriptions rather than the paper itself, such as the exact “19 percentage points” figure in The Jagged Frontier. It was a spot-check, not an audit.
The tools
The claim checker, Find your patterns, the checklist, the quick brief, and Use it now are reading aids. The claim checker and Find your patterns match your text against published rules, which you can read on each page. They do not assess a system, and they run in your browser. The Design Rubric and Readiness Review are checklists. They do not calculate a score.
What this guide is not
- It is not a ranking or benchmark of any AI model or vendor.
- It is not legal, regulatory, or compliance advice, though the Readiness Review points to standards that are relevant.
- It is not peer reviewed, and no outside reviewer has signed off on any pattern yet.
How it changes
Each pattern has a version number and a revision history, and site-wide changes are in the changelog. Corrections are made in the open and logged there. If you find a mistake or a better source, every pattern page has a “Report it” link, or you can use the contact page.