Skip to main content

BetaEarly results, in development

Test methodology

The eval suite measures how accessible AI-generated code is:

We track quality, speed, token usage, and more.

This page outlines the technical details of the test suite.

What we measure

The suite measures how accessible a model's generated code is. It covers WCAG 2.2 Level AA, and best-practice criteria that go beyond the standard.

The suite has three types of tests.

Automated

These are deterministic JavaScript tests using axe-core and DOM assertions.

Browser

These use Playwright in a headless Chrome browser. They interact with the components using keystrokes and clicks, like a person does. The assertions check the accessibility tree directly.

Judgement

These use a second AI model to assess things deterministic JavaScript tests can't decide. For example, whether headings and link text are meaningful.

We compare runs of the suite over time, to see how a change affects quality, speed, token usage and other metrics. In particular:

Back to top

Why our suite

Intopia has spent more than ten years writing and reviewing accessible code and has logged tens of thousands of accessibility issues for clients. That work shows us where development teams most often get accessibility wrong. Our acceptance criteria target those failure points first.

Our tests go beyond conformance. Each check is categorised as either a WCAG requirement or a best practice. Best practice checks cover the things we recommend to clients even though WCAG doesn't require them. These checks cover barriers that users face that aren't covered by WCAG. In our audits we report these as expert observations.

For example, WCAG requires that focus is managed logically when a modal opens. Our best practice check goes further and asks whether focus moves to the modal's heading. For the modal described in our prompt, that's what makes the experience work well for a screen reader user.

We test at two levels. First, isolated components. For example, a form, a modal or a menu button. This shows what a model produces when there's nothing else to think about.

Second, common page templates that combine components, such as a contact page with validation, a landing page with images, a product description page with size and colour selectors, or a project management Kanban board with drag and drop. Each template includes several complex components, so there are many places the model can go wrong. This shows whether the model still gets accessibility right as the task grows.

We push what automation can check. The suite tests keyboard operability, focus visibility and styling, form error UX and viewport reflow. These sit at the edge of what automated testing can do today. For checks that need judgement, we use an LLM as judge.

We check the checker. Automated results aren't taken on trust. We spot-check verdicts (pass, fail, blocked and so on) against the code to confirm they're accurate. We also manually spot-check generated outputs to confirm they work as delivered.

Back to top

Limitations

The suite measures how well a model handles accessibility when it generates code from a fixed set of prompts. These are the main limits on what the results can tell you.

  • Coverage is partial. The suite doesn't cover every WCAG Success Criterion. A component or template that passes every test could still have WCAG failures.
  • No severity ratings. We don't rate the severity of WCAG failures because severity depends on context the suite doesn't have.
  • No assistive technology testing. The suite reads the accessibility tree, which is what assistive technology consumes, but it doesn't run a screen reader or other assistive technology. Real assistive technology testing often finds problems the tree alone doesn't show.
  • Generation only. We test code generated from a prompt. We don't test how a model edits existing code.
  • Fixed prompts. Each generation uses the same prompt. The wording is chosen to avoid steering the model towards accessible outcomes or specific techniques.
  • LLM-as-Judge is non-deterministic. The same generation can get a different grade on a second run of the judge.
  • Cost is notional. The cost figure is for comparison between models only.
  • Three runs per case. More runs would give more certainty at more cost. Three is our current balance.
  • Still expanding. The suite is a snapshot of what we test today, not a complete assessment of a model's ability to produce accessible code.

WCAG Success Criteria coverage

We have at least one test for 27 of the 56 WCAG 2.2 Level A and AA Success Criteria. Each test targets a frequent or notable way to fail that criterion. Passing all our tests for a criterion means the code didn't fail our specific tests, not that it meets the criterion in every case.

Most tests sit under the Perceivable and Operable principles, where common failures are easier to test automatically.

We also have 76 best practice tests that go beyond WCAG.

There are 28 Success Criteria with no tests yet. Some only make sense at page level rather than component level. We're adding page and template tier tests.

Back to top

How the testing works

Each component or template is a case. Each case has a prompt and a rubric.

The prompt is the task we give the model. It's fixed, and worded to avoid steering the model towards accessible outcomes or specific techniques. For example:

Build a Button component in a single self-contained HTML file (inline CSS and JS, no build step, no external libraries). Include two buttons on the page: one enabled button labelled "Save changes" and one disabled button labelled "Delete". Output the complete code in one HTML code block.

The rubric is the list of criteria we grade the output against. Each criterion is either a WCAG Success Criterion or a best practice, and is graded by one of three methods: an automated check, a manual check, or LLM-as-Judge. A rubric changes only when WCAG or our best practice criteria change.

For each case, the model generates code in response to the prompt. A test runner then grades the code against every criterion in the rubric.

Each criterion gets one of six verdicts:

Pass or fail
The check ran and reached a result.
Blocked
The check depends on something that isn't there. For example, a check on a dialog's accessible name is blocked if there's no dialog role to test.
Not applicable
The check is conditional and the condition wasn't met. For example, if the code uses the browser's default focus style, we don't test the contrast of a custom focus style.
Unmeasured
The test harness can't measure this construct.
Error
The check itself failed to run.

Back to top

Harness

The AI SDK harness lets us compare AI models from different vendors under the same conditions. It also measures how much the Intopia web accessibility skill improves each model's code.

How it works

The harness is one agent loop, built on the Vercel AI SDK. It is the same for every model. Each model gets:

  • The same instructions. A short, neutral system prompt. It describes the tools and the rules: work in this folder, do not use the network, and give the full code in the final message.
  • The same tools. Read, Glob, Grep, Write, Edit and RunNode. Each tool can only get to files in a temporary working folder. RunNode runs Node.js scripts from that folder, with no shell and no network access.
  • The same limits. A maximum of 40 steps, a fixed output limit for each step, and a fixed timeout.
  • No sampling settings. We do not send temperature or similar settings. Each model uses its vendor's defaults.

The model gets the task prompt and the tool results, and nothing more.

Measuring the skill's effect

We run each model two times on the same tasks:

  • Skill on. The skill's instructions go into the system prompt.
  • Skill off. The same prompt, without the skill.

The skill is the only difference between the two runs. The change in the scores is the skill's effect on that model.

What we grade

The harness only creates the code. We grade the code with the same process for every model: automated checks with axe-core, browser tests in Playwright that act like a manual tester (keyboard, focus, forced colours), and an independent AI judge for criteria that need human-like decisions. Each result is graded against the WCAG 2.2 AA criteria for that component.

Supported models

The harness connects to models from Anthropic, OpenAI, Google, xAI, Meta and DeepSeek.

Limits of the comparison

The results show how well each model performs in this harness. A model can perform differently in its own vendor's tools.

Back to top

Criteria & rubrics

Each case has a rubric.json: a frozen list of criteria. Here's one criterion from the button case.

{
  "id": "focus-contrast",
  "text": "If the button uses a custom focus style, the focus indicator has a contrast ratio of at least 3:1 against its background colour.",
  "wcagSC": "1.4.11",
  "type": "wcag",
  "method": "manual",
  "check": "focus-contrast",
  "group": "Visual design",
  "dependsOn": ["focus-visible", "keyboard-focusable", "role-button"],
  "condition": "A custom focus style is used - the default browser focus ring is exempt from the 3:1 requirement"
}
text
is the criterion in plain language. It's what a reviewer reads when checking a verdict.
wcagSC
the WCAG success criterion number, or null for a best-practice criterion
type
wcag or best-practice
method
says how it's graded. This one is manual. The other values are automated and judgement (LLM-as-Judge).
check
names the function, or the judge guidance, that grades it.
group
the acceptance criteria grouping: Page structure, Labels and messaging, Semantic markup, Keyboard, Media alternatives, Visual design, or Adaptive UI
dependsOn
(optional) lists the criteria this one needs. If the button isn't keyboard focusable, has no visible focus style, or doesn't have the button role, there's no focus indicator to measure, so this criterion is blocked rather than failed.
condition
(optional) is the rule that decides whether the criterion applies. If the generated button uses the browser's default focus ring, this criterion is not applicable.

A rubric changes only when WCAG or our best-practice criteria change.

Back to top

Component set

The component tier tests isolate one component per case. Acceptance criteria are grouped into the following categories:

  • Label and messaging. Controls have clear names. Forms give instructions, mark required fields and report errors.
  • Semantic markup. Interactive elements expose the correct name, role and state to assistive technology.
  • Keyboard. Everything works with a keyboard. Focus order is logical, focus is always visible, and focus moves correctly when dialogs, menus or new content open.
  • Visual design. Text and controls meet contrast ratios, colour isn't the only signal for meaning, and touch targets are large enough.
  • Adaptive UI. Content reflows at 400% zoom, text resizes and respaces without content being lost.
  • Media alternatives. Images have alt text that matches their purpose. Decorative images are hidden from assistive technology. Icon controls have names.
  • Pointer interaction. Actions don't rely on complex gestures or fire on pointer-down alone. Drag and multi-point gestures have a single-pointer alternative.

The table shows the components and the number of acceptance criteria we test in each category. Darker cells have more criteria. The counts come straight from the test cases, so they stay up to date as criteria are added or removed.

Acceptance criteria tested per component, by category
ComponentLabels and messagingSemantic markupKeyboardMedia alternativesVisual designPointer interactionAdaptive UITotal
Accordion076020116
Button154020012
Card21481101137
Carousel1146060128
Checkbox154030013
Checkbox group274030016
Combobox with list autocomplete31113040132
Combobox without autocomplete11212040130
Disclosure166020116
Icon button13401009
Link254040116
Menu button11210020126
Modal dialog138030116
Radio275030017
Select1107040022
Table14002018
Table with irregular headers160020110
Tabs1118040125
Text field194040018
Toggle button164040015
Toggletip169030120
Tooltip053020111

0 criteria14 criteria

Back to top

Acceptance criteria

These are the acceptance criteria for the button case.

Acceptance CriteriaConformanceTypeCategory
The button has a visible label.WCAG 3.3.2AutomatedLabels and messaging
The button's role is included in the accessibility tree.WCAG 4.1.2AutomatedSemantic markup
The button's accessible name is included in the accessibility tree.WCAG 4.1.2AutomatedSemantic markup
The button's accessible name contains the text of its visible label.WCAG 2.5.3AutomatedSemantic markup
The button's accessible name matches its visible label exactly, or at least starts with the exact text of the visible label.Best practiceAutomatedSemantic markup
If the button is disabled, the button's disabled state is included in the accessibility tree.WCAG 4.1.2AutomatedSemantic markup
The button is focusable using a keyboard.WCAG 2.1.1Browser automationKeyboard
The button is activated by pressing the Space or Enter keys on the keyboard.Best practiceBrowser automationKeyboard
The button has a focus style when it receives focus using a keyboard.WCAG 2.4.7Browser automationKeyboard
The keyboard focus style still renders in forced-colours mode (e.g. Windows High Contrast Mode).Best practiceBrowser automationKeyboard
All text meets the minimum contrast ratio of 4.5:1 against the background colour, or 3:1 for large-scale text.WCAG 1.4.3AutomatedVisual design
If the button uses a custom focus style, the focus indicator has a contrast ratio of at least 3:1 against its background colour.WCAG 1.4.11Browser automationVisual design

Back to top

Future plans

We plan to expand to test:

  • against multiple skills
  • on existing code, not just generated code
  • with other models, including open weight models
  • with frameworks, not just HTML

Back to top