Skip to main content

About

AI models now write a lot of front-end code. We wanted to know how accessible that code is, and where it goes wrong. We built this site to share what we find as we test each model.

In 2026 we ran two long rounds of manual testing while we developed the Intopia web accessibility skill and the Figma Make skill. We then wanted a repeatable set of tests. The tests let us compare models quickly and fairly, and track whether their output changes over time.

We built an automated test suite. With human review, it lets us check specific outputs from selected prompts, benchmark each new model on release, and measure the effect of changes to our skills.

Why components first

Components are the basic parts of every interface. Buttons, tabs, modal dialogs, comboboxes and carousels appear on most sites. If a model gets these wrong, everything built from them inherits the problem.

We test 22 components. We prompt the model to build each one three times, because AI-generated code varies from run to run.

How we test

For each component, we wrote acceptance criteria. They cover the issues we see most often in our audit work, and the best practice techniques we recommend to clients. For example:

  • Can you reach it with a keyboard?
  • Does it have an accessible name?
  • Does it tell screen reader users what state it's in?
  • Does focus go to the right place?

We check every output against these criteria. That's more than 1,200 checks per model. A person reviews every result before we publish it.

What's next

We'll test new models as they're released and publish what we find here. We'll keep developing skills to fill gaps in model knowledge and correct common mistakes. We're also building template tests, which ask models to build complete layouts instead of single components.