Two blocks of code, produced from nearly the same three-word prompt to a coding assistant, show exactly what vibecoding, the practice of generating a working component or page from a natural-language description, tends to get wrong.
Submit
Both versions render identically in a browser. A sighted mouse user cannot tell them apart, and that is precisely the problem: the first block passes every review a team is likely to run, while carrying none of the structure that assistive technology and keyboard navigation depend on. The second block is not more sophisticated code, nor is it meaningfully longer. It simply uses the elements the web platform already provides for exactly this purpose, a fix accessibility practitioners have been recommending since long before a model could write a line of code on its own.
Why the gap exists
A large language model (LLM) generates code by predicting the statistically likely continuation of a pattern it has seen many times, and the training corpus for front-end code is saturated with div-based layouts, placeholder-only inputs, and icon buttons with no accessible name. Ask for "a clean modern signup form" and a coding assistant will often return exactly that: a form satisfying every visual convention of recent web design while failing several of the oldest rules in the Web Content Accessibility Guidelines. This is not a defect unique to one vendor's model; it reflects the shape of the code those models were trained on, most of which was never checked for accessibility in the first place. A generated component inherits the accessibility debt of the entire web, compressed into a single confident-looking output.
Four ways it breaks
The failures are not exotic. They are the same handful of problems accessibility auditors have flagged for two decades, reappearing at the speed of autocomplete.
| Failure | What ships | What's missing | The fix |
|---|---|---|---|
| Missing or mismatched labels | A placeholder attribute standing in for a label, or a label paired with the wrong control through a missing or incorrect for attribute. | An accessible name a screen reader can announce once the user starts typing and the placeholder disappears. | A real element programmatically associated with its input. |
| Broken focus order | A page assembled from several independently generated sections, each internally consistent but never checked against the others. | Agreement between the document object model order and the visual order CSS produces. | A keyboard pass through the page, not an assumption drawn from the layout. |
| Div-based controls | A | Default keyboard support, a role, and state that assistive technology can announce. | A native , or an Accessible Rich Internet Applications (ARIA) role applied only where no native element solves the problem. |
| Silent state changes | Modals, dropdowns, and validation messages that appear and disappear in the DOM without any accompanying signal. | A cue a screen reader user can hear when the state actually changes. | An ARIA live region or focus management tied directly to the state change. |
Individually, each of these is a known, well-documented pattern with an established fix; none of them is exotic, and accessibility auditors have been flagging the same handful of problems for two decades. What changes with vibecoding is not the nature of the mistakes but their volume and speed: a single prompt can now generate dozens of these failures across an entire page in seconds, and a team under deadline pressure has every incentive to accept code that already looks finished.
Why visual review doesn't catch it
Visual review is the review most teams actually perform, and generated interfaces are specifically optimized to pass it, because a pull request that renders correctly in a browser preview clears the bar most reviewers apply by default. The reviewer is judging a screenshot in their head, not reading the markup underneath it. Automated linting catches some of this territory, a missing alt attribute, for instance, but a div standing in for a button will not trip most build pipelines, since nothing in the syntax is technically wrong. As Signal & Syntax explored in how AI agent browsers read and fill out web forms, an agent parsing that same placeholder-only field runs into a version of the identical gap a screen reader hits: a control that looks complete and communicates nothing about what it actually is.
I think the deeper problem is cultural rather than technical: teams have started treating generated code as a finished draft rather than a first one, and the confidence of the tool's output makes that mistake easy to make. A junior developer's pull request gets read with some suspicion; a generated component that compiles cleanly and matches the design mock tends to get merged with less scrutiny, precisely because nobody wrote it, so nobody feels responsible for its assumptions.
Where the review has to sit
None of this argues against using AI pair-programming tools. It argues for placing the accessibility review at a specific point in the workflow rather than treating it as an optional pass that happens if time allows. The most effective placement is before a generated component is styled, not after: reviewing markup in its raw, unstyled state forces attention onto structure, whether the form field carries a real label, whether the interactive element uses a native element or an appropriate ARIA role, whether the document order makes sense when read top to bottom without the layout to lean on. The ARIA Authoring Practices Guide is the standing reference for this kind of check, since it documents the expected keyboard behavior and role for every common interactive pattern, including the cases where a native HTML element already solves the problem and ARIA should not be reached for at all. WebAIM's guidance on keyboard accessibility testing covers the focus-order check specifically, and it requires nothing more than a keyboard and a few minutes per component.
Practically, this means adding one gate to the pull request template that a linter cannot fill in: a reviewer confirms, by reading the diff, that every interactive element is a native control or carries the correct role and keyboard behavior, that every form field has a programmatically associated label, and that focus order was checked with a keyboard rather than assumed from the layout. That single step catches most of what vibecoded components get wrong, and it costs a few minutes per component rather than a redesign. It belongs on the same checklist as the piece on cookie banners becoming an API surface for AI agents: both are cases of markup doing double duty as something more than decoration, read by a consumer other than a sighted mouse user.
Vibecoding did not invent inaccessible markup, and it will not be fixed by asking a model to try harder on a single prompt. It has made the old failures cheaper to produce and easier to miss, which means the review step that used to feel optional now has to be treated as load-bearing. The tools are fast. The judgment about whether their output actually works for every user still has to be human, and it still has to happen before the component ships, not after a complaint arrives.