Why a 95% pass rate isn't an accessible website
Most models pass 88% to 97% of our acceptance criteria across 22 common components. But the failures that remain can create barriers for people with disabilities, and they add up quickly across a page.

With AI models writing more and more front-end code, we wanted to know how accessible that code is, and where it goes wrong. We built this site to share what we find as we test each model. It shows which of our acceptance criteria each model passes and fails when it builds common components.
A starting point for formal testing
After two lengthy rounds of manual testing in 2026 while developing the Intopia web accessibility skill and the Figma Make skill, we wanted a repeatable set of tests to compare models quickly and fairly and to track whether their outputs’ accessibility changes over time. So we developed an automated test suite which, combined with human review, allows us to check specific outputs from selected prompts, benchmark each new model on release and measure the effectiveness of changes that we make to our skills.
Why components first
We started with components because they're the basic building blocks of every interface. Buttons, tabs, modal dialogs, comboboxes and carousels appear on most sites. If a model gets these wrong, everything built from them inherits the problem.
At the time of writing, we test 22 components, prompting the model to build each component three times to cover a range of outputs, since AI-generated code is inherently variable.
Acceptance criteria based on experience and best practice
For each component, we have written acceptance criteria that cover the most common accessibility issues we see in our work as accessibility professionals, plus the best practice techniques we recommend to clients.
Can you reach it with a keyboard? Does it have an accessible name? Does it tell screen reader users what state it's in? Does focus go to the right place?
We check each model's output against this criteria. At three generations per component, that's more than 1,200 checks per model! A person reviews every result before we publish it.
What the results show
On the surface, the results seem impressive. Most models pass between 88% and 97% of our acceptance criteria:
- Claude Fable 5.1, Claude Opus 5.5 and GPT-6 Astra: about 97%
- Grok 4.7, GPT-6.1 Sol, GPT-6 Luna, Claude Sonnet 5.5: 88% to 96%
However lots of passing tests doesn’t mean the code is highly accessible. When we audit websites, we report what fails, not what passes.
When we look at the failing tests for each component, we notice that models in the 88% to 96% range have about two failing tests for each component. This adds up quickly when looking at a whole page and a whole site.
We often see this in audits: one bad component repeated across a page can have a serious impact on the accessibility of the page as a whole. If each component has one or two issues, the number of reported problems compounds.
Even the best models fail on average once every two components.
What's next
We'll test new models as they're released and publish what we find here.
We’ll keep developing skills to fill gaps in the model’s knowledge and correct common mistakes.
We're also working on template tests that will challenge models to build complete layouts, not single parts.
