The Number That Should Bother You

Of the 55 success criteria in the Web Content Accessibility Guidelines (WCAG) 2.2 AA, the standard that defines whether a website is accessible to people with disabilities, automated tools can reliably evaluate seven.

Seven out of fifty-five. That is 13%.

The other 48 criteria require something machines cannot do yet: understand intent, evaluate meaning, judge whether a user experience actually works for a human navigating by ear or by keyboard alone. A tool can verify that an image has an alt attribute. It cannot verify that the alt text accurately describes the image. A tool can confirm a button exists. It cannot confirm the button makes sense out of visual context.

This is the automation boundary, and it is the single most important thing to understand before making any investment decision about AI and web accessibility.

AI Is a Power Tool, Not an Autopilot

Here is the direct answer to whether AI can handle your accessibility compliance: it cannot. The reason is not that the technology is bad, but that the problem is structurally incompatible with full automation.

Web accessibility under WCAG, a World Wide Web Consortium (W3C) standard organized around four principles (Perceivable, Operable, Understandable, Robust), is a mix of mechanical rules and human judgments. Some criteria are binary and machine-verifiable: does this image have alt text? Is this color contrast ratio above 4.5:1? Does this page have a language attribute? Automated tools handle these reliably.

The majority of WCAG criteria, however, require evaluating meaning, context, or user experience. Is this alt text accurate? Does this form's error message help the user fix the problem? Can a keyboard-only user actually complete this workflow? These are judgment calls, and they remain judgment calls whether the tool evaluating them uses rule-based pattern matching or a large language model.

The automation boundary is not a temporary limitation waiting for better models. It is a structural property of the standard itself. WCAG measures whether humans can use websites. Machines can check the scaffolding; they cannot check the experience.

None of this means AI is useless for accessibility. It means AI is a power tool operating in a specific range, and understanding that range (where it is reliable, where it is partial, where it is blind) is how you deploy it intelligently instead of deploying it as a substitute for the work it cannot do.

Three Numbers That Are All Correct

You will encounter three different claims about how much of web accessibility automated tools can catch. All three are real, sourced, and defensible. They answer different questions.

13%: Criteria Reliability

Accessible.org audited all 55 WCAG 2.2 AA success criteria against the capabilities of current automated tools (as of 2025). Result: 7 criteria (13%) can be reliably flagged by automation. Another 25 criteria (45%) are partially detectable, meaning the tool can catch some violations but not all instances. The remaining 23 criteria (42%) are invisible to automated tools entirely. This measures how many rules a machine can fully enforce. It is the strictest frame, and the most useful for understanding structural limits.

31%: Criteria With Any Automated Rules

The W3C ACT (Accessibility Conformance Testing) Task Force maintains a set of approved automated testing rules. As of Adrian Roselli's April 2025 analysis, approved rules exist for 31% of WCAG 2.2 Level A and AA criteria. This is higher than 13% because it includes partial coverage: criteria where a rule can detect some violations even if it cannot detect all of them. It measures how many criteria have at least one automated check, not how many are fully covered.

57%: Issue Volume

Deque's 2021 study of over 13,000 pages and nearly 300,000 individual accessibility issues found that automated testing with axe-core caught 57.38% of issues by volume. This is the most favorable number for automation, and it is legitimately measured, but it reflects a different question: of the issues that actually occur on real websites, how many are the kinds that tools can detect? The answer is skewed by frequency. Missing alt text, insufficient color contrast, empty links, and missing form labels are both machine-detectable and extremely common. The WebAIM Million 2026 report found that just six error types account for 96% of all detected homepage failures. Automated tools are good at catching the errors that happen the most.

13%

Criteria automation can fully enforce

7 of 55 WCAG 2.2 AA criteria are reliably flagged. The strictest frame, measuring complete enforcement.

Accessible.org, 2025

31%

Criteria with at least one automated rule

Higher because it counts partial coverage: a rule that catches some violations still counts.

W3C ACT Task Force, via Roselli, Apr 2025

57%

Real-world issues caught, by volume

Skewed by frequency: the most common error types are also the most machine-detectable.

Deque, Mar 2021, 13,000+ pages

Figure 1. Three statistics about automated accessibility coverage. All three are correct; each answers a different question, which is why they appear to contradict each other.

So which number is right? All three. The 57% tells you that automated tools are worth running, because they catch more than half of real-world issues by count. The 31% tells you that the standards body has only managed to write reliable automated checks for less than a third of the criteria. The 13% tells you that the floor of truly reliable, comprehensive machine coverage is low. As for the 42% that are invisible to any automated tool, those are the criteria where human judgment is not optional; it is the only option.

13%
45%
42%

Reliable7 criteria

Color contrast (1.4.3). The luminance ratio between text and background is deterministic math, so a tool computes it exactly.

Partial25 criteria

Heading structure (1.3.1). A tool checks that heading levels run in order; it cannot check whether the headings describe their sections.

Invisible23 criteria

Consistent navigation (3.2.3). Judging whether navigation is "the same" across differing layouts needs context a DOM scan cannot reconstruct.

Figure 2. The automation boundary across all 55 WCAG 2.2 AA success criteria, with one representative criterion per tier.

What Automation Actually Evaluates

To understand the automation boundary concretely, it helps to walk through what WCAG criteria look like at each level of detectability.

Reliably Detectable (the 13%)

These are the criteria with clear, binary, machine-checkable properties.

  • Image alt text presence (WCAG 1.1.1, partial): A tool scans the Document Object Model (DOM), the browser's live representation of the page's HTML structure, and checks whether <img> elements have alt attributes. Present or absent. Machine-readable. Note, though, that this checks presence, not quality. The same criterion has a reliable-detection component (does the attribute exist?) and a human-judgment component (is the description accurate?). Most "reliably detectable" criteria have this split.
  • Color contrast ratios (WCAG 1.4.3): A tool computes the luminance ratio between foreground text and its background. The math is deterministic, 4.5:1 for normal text and 3:1 for large text. Tools handle this well, with caveats for text over images or gradients.
  • Page language declaration (WCAG 3.1.1): Is there a lang attribute on the <html> element? Yes or no.

Partially Detectable (the 45%)

These criteria have machine-checkable components wrapped around human-judgment cores.

  • Heading structure (WCAG 1.3.1): A tool can check whether headings follow a logical numeric sequence (h1, h2, h3, with no skipped levels). It cannot evaluate whether the headings accurately describe their sections. A page with perfectly ordered headings that say "Section 1," "Section 2," "Section 3" passes the automated check and fails the accessibility intent.
  • Link purpose (WCAG 2.4.4): A tool can flag links whose visible text is "click here" or "read more." It cannot evaluate whether a link with descriptive text actually describes the right destination.
  • Error identification (WCAG 3.3.1): A tool can detect whether a form has error messages. It cannot evaluate whether those error messages help the user understand what went wrong and how to fix it.

Invisible to Automation (the 42%)

These criteria depend entirely on context, meaning, or interaction patterns that machines cannot evaluate.

  • Audio descriptions and captions accuracy (WCAG 1.2.x): A tool can check whether a <track> element exists on a video. It cannot evaluate whether the captions accurately represent the audio content, whether audio descriptions adequately convey visual information, or whether the timing is synchronized with the content it describes.
  • Consistent navigation (WCAG 3.2.3): Does the navigation appear in the same relative order across pages? Answering that requires understanding what "same navigation" means across different page layouts, a judgment that depends on visual and structural context a DOM scan cannot reconstruct.
  • Content on hover or focus (WCAG 1.4.13): When content appears on hover or focus, can the user dismiss it, move the pointer over it, and does it persist until dismissed? Testing this requires simulating interaction sequences and evaluating whether the resulting behavior meets user expectations, which is not a static DOM property.

This is the mechanism behind the numbers. The 42% is not a gap that better algorithms will close; it is a category of criteria that require understanding what content means to a human user. Until a machine can do that reliably, these criteria require human evaluation.

The Overlay Experiment, and What It Proved

If you cannot automate accessibility fully, can you automate it partially and then bolt on a fix layer for the rest? That is the premise behind accessibility overlays, third-party JavaScript widgets that inject themselves into websites and attempt to repair accessibility issues at runtime.

The overlay market has been running this experiment at scale for several years. The results are in.

The Federal Trade Commission (FTC) Intervened

In January 2025, the FTC filed an enforcement action against accessiBe, the largest overlay vendor, for deceptive marketing claims. AccessiBe had marketed its AI-powered overlay as capable of making "any website WCAG compliant." The FTC found these claims unsupported by evidence. The final order, issued in April 2025, imposed a $1 million fine and prohibited specific compliance claims (Lainey Feingold's analysis covers the settlement terms in detail). This was not a competitor complaint or an advocacy campaign; it was a federal agency concluding that the core marketing premise of AI overlay compliance was deceptive.

The Lawsuit Data Contradicts the Value Proposition

UsableNet's litigation tracker recorded 3,117 federal web accessibility lawsuits in 2025, a 27% increase over 2024. Of those lawsuits, overlays appeared on the defendant's site in 22.6% of cases during the first half of 2025. If overlays delivered compliance, this number should be zero or near it. It is not. Sites with overlays installed get sued at significant rates, which suggests the overlay is not providing the legal protection it is marketed to deliver.

The Research Measured What Overlays Fix

A peer-reviewed study published at the Association for Computing Machinery Special Interest Group on Accessible Computing (ACM SIGACCESS) conference in 2025 evaluated AI-powered overlay effectiveness and found overlays correct only 20% to 40% of common accessibility issues, missing core requirements like semantic heading structure, form labeling, and keyboard navigation. A separate ACM SIGACCESS 2024 study examined the experience of blind and low-vision users with overlays directly. Participants reported that overlays made sites less usable: overlay-provided screen readers conflicted with users' own assistive technology, intercepted keyboard commands, and reordered page elements unpredictably. Users described opting to avoid overlay-equipped sites entirely.

One caveat worth stating plainly: both ACM papers were paywalled during research for this article. The findings described here are drawn from published summaries and are consistent with broader evidence, but specific methodology details have not been independently verified against the full papers.

The Professional Community Opposes Them

The European Disability Forum and the International Association of Accessibility Professionals issued a joint statement in May 2023 that overlays "do not constitute an acceptable alternative or substitute for fixing the website itself." The Overlay Fact Sheet, an open statement opposing overlays, has been signed by over 1,000 accessibility practitioners and disability advocates, including WCAG and Accessible Rich Internet Applications (ARIA) specification contributors and accessibility engineers at Google, Microsoft, Apple, Shopify, and the BBC.

The overlay model fails because it tries to fix the symptoms (DOM-level ARIA attributes, contrast adjustments, widget toolbars) without addressing the cause (inaccessible design and code). It operates on the 13% of criteria that are machine-detectable, applies partial fixes to some of the 45% that are partially detectable, and cannot touch the 42% that require human judgment. The runtime injection approach also introduces its own accessibility problems, conflicting with assistive technology, altering expected page behavior, and creating a dependency on third-party JavaScript that can break with any site update.

Where AI Actually Helps

The overlay failure does not indict AI for accessibility. It indicts the specific model of bolting AI onto a broken site and calling it compliant. Within its automation boundary, AI does useful work.

Automated Scanning in CI/CD

Tools like axe-core, integrated into continuous integration and continuous deployment (CI/CD) pipelines, catch the detectable 57% of issues before code ships. This is valuable. Catching a missing alt attribute or a contrast violation in a pull request is cheaper and faster than finding it in a quarterly audit. The key is treating the scan as a smoke alarm rather than a building inspector: it catches the common fires, but it does not certify the building.

AI-Generated Alt Text for Simple Images

The AltGen research paper (2025) found AI-generated alt text achieved 0.93 cosine similarity with human-written reference descriptions and a 0.76 Bilingual Evaluation Understudy (BLEU) score, both standard measures of how closely generated text matches human-written references, with a 97.5% reduction in accessibility errors in Electronic Publication (EPUB) documents. For photographs with clear subjects (a person, a product, a landscape) AI-generated descriptions are approaching useful quality. The boundary sharpens for complex images: charts, diagrams, infographics, and images where the communicative intent depends on context (a photo used as evidence versus the same photo used decoratively). How well AI handles these cases is an open question with limited published evidence.

Pattern Detection at Scale

AI can scan large codebases for common anti-patterns, including missing form labels, empty buttons, and improper heading hierarchies, faster than manual review. Roselli's April 2025 analysis found that manual review catches 7.5 times more issues than the best automated tool, but the automated tool catches its issues in seconds across thousands of pages. The combination of automated breadth and manual depth is more effective than either alone.

Assistive Technology Itself

Screen readers, chiefly JAWS (40.5% primary desktop usage), NVDA (37.7%), and VoiceOver (dominant on iOS) according to WebAIM's Screen Reader User Survey #10 (December 2023 to January 2024, 1,539 respondents), are sophisticated software that uses AI-adjacent techniques to interpret the accessibility tree.

The accessibility tree is the browser's representation of a page's structure as exposed to assistive technologies: a filtered version of the DOM containing only accessible objects with their roles, names, states, and relationships. Screen readers navigate this tree, not the visual layout. When the tree accurately represents the page, screen readers work. When it does not, no amount of overlay JavaScript can fix the mismatch.

The pattern is consistent: AI works well when the task is narrow, well-defined, and has clear success criteria that a machine can verify. Detect missing alt text, yes. Generate alt text for a simple photo, increasingly yes. Evaluate whether a complex multi-step form workflow is usable by a screen reader user, no.

The Regulatory Timeline Engineering Leaders Need

If the technical argument is that AI covers part of the problem, the regulatory argument is that the law requires all of it. Three developments in 2024 and 2025 created a compliance timeline that engineering leaders need on their roadmaps.

The Department of Justice (DOJ) Codified WCAG as Law (April 2024)

The DOJ finalized a rule under Americans with Disabilities Act (ADA) Title II requiring state and local government websites and mobile applications to conform to WCAG 2.1 Level AA. This is the first federal rule that explicitly adopts WCAG as the legal standard for state and local government digital properties, distinct from Section 508 of the Rehabilitation Act, which applies to federal agencies. The original rule set compliance deadlines of April 2026 and April 2027; an April 2026 interim final rule extended these to April 2027 for entities serving populations of 50,000 or more, and April 2028 for smaller entities and special districts.

The European Accessibility Act Took Effect (June 2025)

The EAA requires digital products and services across the EU to be accessible, referencing the EN 301 549 standard, which maps to WCAG 2.1 Level AA. It applies to e-commerce, banking, telecom, transport, and other consumer-facing digital services. Organizations operating in the EU market face compliance obligations now, not at some future deadline.

The FTC Enforced Against Deceptive AI Claims (April 2025)

The accessiBe enforcement action signals that regulators are not only setting accessibility standards, they are actively penalizing companies that market AI as a shortcut around those standards. The $1 million fine and the specific prohibition on compliance claims establishes a precedent: claiming AI makes your site WCAG-compliant when it does not is a deceptive trade practice.

What This Means for Your Roadmap

The regulatory trajectory points in one direction: WCAG compliance as a legal requirement, with no carve-out for automated or AI-based alternatives. If your organization is subject to ADA Title II (state and local government), the EAA (EU market), or Section 508 (federal contractors), the compliance clock is running. If your organization faces ADA Title III exposure (private sector, where most of the 3,117 federal lawsuits in 2025 were filed), the litigation risk is already present and growing at 27% year over year.

The investment question is not whether AI can make this cheaper. It is what combination of tools and processes actually achieves compliance. AI-based scanning tools are part of that combination; they are not the whole thing, and marketing them as the whole thing is now explicitly a regulatory risk.

Why the Problem Is Getting Worse

Here is what should alarm everyone building for the web: despite the wide availability of accessibility tools, accessibility is getting worse.

The WebAIM Million 2026 report, an annual automated evaluation of the top one million homepages, found 95.9% of pages had detectable WCAG failures, with an average of 56.1 errors per page. Errors per page increased 10.1% from the prior year, reversing six consecutive years of gradual improvement.

Consider what this means. These are detectable errors, the 13% of criteria that automated tools can reliably evaluate, which is to say the simplest and most machine-checkable accessibility failures. They are increasing on the most-visited websites in the world, during a period when AI-powered accessibility tools have never been more widely available.

The most likely explanation is not that the tools are failing. It is that the tools are being treated as sufficient. Organizations deploy an overlay or run an automated scan, see a compliance score, and move on, without addressing the 87% of criteria that the scan cannot evaluate and, apparently, without fixing all of the 13% it can.

This is the real cost of the "AI is your accessibility department" narrative. It does not just leave the hard problems unsolved; it creates a false sense of completion that discourages solving the easy ones.

What Actually Works (and Why)

If the automation boundary explains what machines cannot do, the remaining question is what fills the gap. Three practices, used together, address the full spectrum of WCAG criteria. None of them is new, and all of them require organizational commitment rather than just tooling.

Shift-Left Accessibility

Shift-left accessibility means integrating accessibility into the earliest stages of design and development, before there is code to scan. This includes accessible design systems with pre-built components that carry correct ARIA semantics (roles, states, and properties like aria-label and role="navigation"). It includes developer training on semantic HTML and keyboard interaction patterns. It includes accessibility acceptance criteria in user stories and design review checkpoints that evaluate contrast, focus order, and content structure before implementation begins.

In terms of the automation boundary, shift-left addresses the 42% of undetectable criteria at the design stage, where fixing them is cheapest. A form designed with clear labels, logical tab order, and helpful error messages from the start does not need a machine to retrofit those qualities later. The accessible design decision is made by a human who understands the user's experience, which is exactly what the undetectable criteria measure.

Manual Audit Paired With Automated Scanning

Manual audit paired with automated scanning combines the breadth of automation (fast coverage of the detectable criteria across every page) with the depth of expert review (evaluation of meaning, context, and interaction for the criteria automation cannot touch). Roselli's finding that manual review catches 7.5 times more issues than the best automated tool is not an argument against automated tools; it is an argument for combining them. Run axe-core or WAVE across every page, and have an accessibility specialist review the critical user journeys. The automated scan sets the baseline; the manual audit finds what the baseline cannot.

Inclusive User Testing

Inclusive user testing means testing with people who actually use assistive technology. The WebAIM Screen Reader User Survey #10 found that about two-thirds of screen reader users never or rarely contact site owners about accessibility barriers. The barriers are there; the feedback is not arriving. Organizations that test with screen reader users, keyboard-only users, users with motor impairments, and users with cognitive disabilities discover the failures that neither automated tools nor sighted accessibility auditors reliably catch.

These three practices are not a checklist to implement once. They are ongoing processes, in the same way that security is not a one-time penetration test but a continuous practice embedded in development. The parallel is not accidental: accessibility, like security, is a property of the entire system rather than a feature you bolt on.

What We Do Not Know Yet

Several questions remain open, and intellectual honesty requires naming them.

WCAG 3.0 and AI-Based Evaluation

WCAG 3.0 is in draft and may fundamentally change how accessibility conformance is measured, potentially introducing outcome-based testing that could expand the automation boundary. The current status and direction are not settled enough to predict.

AI Alt Text for Complex Images

The AltGen research showed promising results for image descriptions broadly, but the quality of AI-generated alt text for complex images (charts, diagrams, infographics) versus simple photographs is not well studied. The distinction matters because complex images are where alt text carries the most information and where getting it wrong is most harmful.

Overlay Versus Remediation Outcomes

No controlled study has been published comparing accessibility outcomes for sites using overlays against sites that underwent manual remediation, holding other variables constant. The evidence against overlays is strong but indirect: FTC enforcement, lawsuit data, user experience studies, and professional community opposition. A direct controlled comparison would be more definitive.

EAA Enforcement Against Overlays

The European Disability Forum opposes overlays, but no enforcement action under the European Accessibility Act has specifically targeted overlay use. Whether EU regulators will follow the FTC's precedent is unknown.

The Net Effect of AI on Accessibility

AI tools detect real issues and generate useful fixes within their automation boundary, while AI marketing convinces organizations they have solved a problem they have not solved. The WebAIM Million 2026 suggests the net effect may currently be negative, since accessibility is getting worse during the period of greatest AI tool availability. Correlation is not causation, however, and multiple factors (increased web complexity, framework-heavy development, reduced accessibility training) could explain the trend independently.

The Boundary Is the Strategy

The automation boundary is not a limitation to work around. It is the map that tells you where to deploy each type of resource.

For the 13% of criteria that machines evaluate reliably, automate aggressively. Put axe-core in your CI/CD pipeline. Fail builds on contrast violations and missing alt text. Let machines do what machines do well, at scale, on every commit.

For the 45% of criteria with partial detectability, use automated tools as triage. They will find some violations in these categories. Treat every automated finding as real, but do not treat the absence of automated findings as clearance. The tool caught what it could see; a human needs to evaluate what it could not.

For the 42% that automation cannot touch, human expertise, inclusive design, and user testing are not optional extras. They are the only path to compliance. No amount of AI advancement changes this until machines can evaluate whether a human experience is coherent and usable, and anyone selling that capability today is selling something that does not exist.

AI is a power tool. Use it where it works, know where it stops, and build the rest yourself.