Tool
AI Design Evaluation Rubric
A quick scoring checklist for judging a single AI feature or design. Each of the six groups has three checks. Use it in a design critique or before a launch review, then follow the checks that fail back to the patterns.
AAgency — surface, don’t decide
- A1The system presents options, evidence, or analysis rather than silently executing irreversible actions on the user’s behalf.
- A2A human can see what the AI is about to do before it happens, not only after.
- A3High-stakes actions (financial, legal, personal data, compliance, employment) always route through explicit human confirmation.
TTransparency — reasoning is legible
- T1A confidence level or certainty signal is shown, not just a final answer.
- T2The reasoning behind a suggestion is visible in plain language, not buried in a tooltip or absent entirely.
- T3A user can trace an output back to the source data or logic that produced it.
HHonesty — limits are disclosed
- H1The system has an explicit “can’t determine this” state, rather than guessing with false confidence.
- H2Edge cases and low-confidence outputs are flagged distinctly from high-confidence ones, not blended together in the same visual treatment.
- H3Product or marketing copy around this feature doesn’t oversell what the AI can reliably do.
EEquity — works across different users
- E1The design was tested with assistive technology (screen reader, keyboard-only, voice control), not just visually reviewed.
- E2The system accounts for different user contexts (expert vs. novice, high-stakes vs. routine) rather than a single path for everyone.
- E3Output has been checked for systematic bias against any user group, language, region, or edge-case identity.
RReal reduction — not hidden complexity
- R1The AI measurably reduces steps, tools, or time, not just visual clutter while the underlying complexity remains.
- R2Automation is reversible, a user can undo, edit, or override without starting the whole task over.
- R3The feature has been validated with real task-timing or usage data, not only a demo or stakeholder walkthrough.
FFailure design — graceful degradation
- F1There’s a clear, designed path for what happens when the AI is wrong, not just a generic error state.
- F2Unfixable or ambiguous cases are flagged explicitly and don’t silently block or corrupt unrelated work.
- F3A human reviewer can always intervene without needing engineering support to do it.
This page is a reference list of the checks: it does not calculate a score. Check numbers (A1, T2, and so on) are labels used on this site.