Tool

AI Design Evaluation Rubric

A quick scoring checklist for judging a single AI feature or design. Each of the six groups has three checks. Use it in a design critique or before a launch review, then follow the checks that fail back to the patterns.

AAgency — surface, don’t decide

  1. A1The system presents options, evidence, or analysis rather than silently executing irreversible actions on the user’s behalf.
  2. A2A human can see what the AI is about to do before it happens, not only after.
  3. A3High-stakes actions (financial, legal, personal data, compliance, employment) always route through explicit human confirmation.

TTransparency — reasoning is legible

  1. T1A confidence level or certainty signal is shown, not just a final answer.
  2. T2The reasoning behind a suggestion is visible in plain language, not buried in a tooltip or absent entirely.
  3. T3A user can trace an output back to the source data or logic that produced it.

HHonesty — limits are disclosed

  1. H1The system has an explicit “can’t determine this” state, rather than guessing with false confidence.
  2. H2Edge cases and low-confidence outputs are flagged distinctly from high-confidence ones, not blended together in the same visual treatment.
  3. H3Product or marketing copy around this feature doesn’t oversell what the AI can reliably do.

EEquity — works across different users

  1. E1The design was tested with assistive technology (screen reader, keyboard-only, voice control), not just visually reviewed.
  2. E2The system accounts for different user contexts (expert vs. novice, high-stakes vs. routine) rather than a single path for everyone.
  3. E3Output has been checked for systematic bias against any user group, language, region, or edge-case identity.

RReal reduction — not hidden complexity

  1. R1The AI measurably reduces steps, tools, or time, not just visual clutter while the underlying complexity remains.
  2. R2Automation is reversible, a user can undo, edit, or override without starting the whole task over.
  3. R3The feature has been validated with real task-timing or usage data, not only a demo or stakeholder walkthrough.

FFailure design — graceful degradation

  1. F1There’s a clear, designed path for what happens when the AI is wrong, not just a generic error state.
  2. F2Unfixable or ambiguous cases are flagged explicitly and don’t silently block or corrupt unrelated work.
  3. F3A human reviewer can always intervene without needing engineering support to do it.

This page is a reference list of the checks: it does not calculate a score. Check numbers (A1, T2, and so on) are labels used on this site.