Open any web page in a browser and you see a designed surface: typography, color, images, layout, whitespace. Open the same page in ChatGPT, Perplexity, or any retrieval-augmented LLM, and it sees something radically different. There is no rendering engine, no pixel grid, no visual hierarchy. The model receives text, and the structure of that text depends almost entirely on decisions you made (or did not make) in your HTML.
This is not an abstraction. It is a concrete, testable difference in what information survives the journey from your server to the model's context window. Understanding what survives, what gets lost, and what gets misinterpreted is the first step toward building pages that AI systems can reliably read.
What actually arrives
When an LLM-powered system reads a web page, it does not render the page the way a browser does. The retrieval pipeline fetches the HTML (or a pre-processed version of it), strips or parses the markup, and converts the result into a text representation that fits inside the model's context window. The details vary by system, though the general pattern resembles the extraction work done by tools such as Mozilla Readability: identify the main content, remove boilerplate, and preserve the text structure that remains.
Step 1
Fetch
Request HTML from server
Step 2
Parse
Build document structure
Step 3
Extract
Separate content from chrome
Step 4
Convert
HTML to text / Markdown
Step 5
Chunk
Split into retrievable units
Figure 1. The five-stage pipeline between your server and an LLM's context window. Step 3 (Extract) is where semantic HTML matters most.
Fetch. The system requests the page's HTML from the server, much like a search engine crawler. JavaScript-rendered content may or may not be available, depending on whether the retrieval system executes JavaScript. Google's JavaScript SEO guidance is useful context here because it separates fetch, render, and index behavior; retrieval systems do not all make the same rendering guarantees. Content that only appears after client-side rendering may therefore be invisible to the model entirely.
Parse. The raw HTML is parsed into a document structure. Semantic HTML elements like <article>, <nav>, <main>, <aside>, and heading tags (<h1> through <h6>) provide structural signals. Class names and IDs are typically discarded or ignored, because they are project-specific strings with no universal meaning.
Extract. This is where semantic HTML matters most. The system identifies the "main content" and separates it from navigation, footers, sidebars, cookie banners, and other boilerplate. An <article> element explicitly marks the primary content, while a <nav> element explicitly marks infrastructure to skip. Without these signals, the extractor falls back on heuristics (text density, element position, common class-name patterns) that are less reliable and more prone to error.
Convert. The extracted content is converted into a text representation that often resembles Markdown: headings become # and ## lines, lists become bulleted or numbered text, code blocks are preserved, and links may be converted to inline references or footnotes. This text is what the model actually receives in its context window.
Chunk. For longer pages, the text is split into chunks (typically 500 to 2,000 tokens) that can be individually embedded and retrieved. Heading structure directly influences how chunking works: a well-structured page with descriptive <h2> headings produces chunks that align with topical boundaries, while a page with no headings or generic headings ("Section 1," "More Info") produces chunks that split at arbitrary points with no regard for the content's logical structure.
The result of this pipeline is that the model sees your content through a narrow keyhole. CSS is gone. Layout is gone. Images are gone unless alt text is present. Color, font weight, and visual emphasis are gone. What remains is text, structure, and the semantic relationships encoded in your HTML.
I spent a weekend feeding the same article to five different retrieval systems with two versions of the markup (semantic versus div-only), and the extraction differences were significant enough to change the model's summary of the page. That is not a theoretical concern; it is a measurable one.
What gets lost
The losses are specific and predictable.
Figure 2. What information survives the retrieval pipeline versus what the model never receives.
Visual hierarchy disappears. A <div class="hero-title"> styled with 48px bold font looks dominant on screen, yet to an HTML parser it is an anonymous container with no structural significance. A proper <h1> element, by contrast, survives the parsing step and appears as the top-level heading in the extracted text regardless of how it is styled. The visual weight you gave an element through CSS does not transfer to the model's understanding.
Layout relationships vanish. A sidebar positioned to the right of the main content, visually separated by whitespace and a border, is not "to the right" in the HTML. It is a block of elements that appears somewhere in the document flow. If it is wrapped in an <aside> element, the parser knows it is supplementary. If it is a <div class="sidebar">, the parser might include it in the main content, interleave it with article text, or discard it unpredictably.
Images become their alt text. When the parser encounters an <img> tag, the only information that survives into the text representation is the alt attribute. An image with alt="Diagram showing how the accessibility tree maps to the DOM" contributes that sentence to the model's context. An image with alt="" or no alt attribute contributes nothing. An image with alt="image1.png" contributes noise.
For content that relies heavily on diagrams, charts, or visual explanations, this is a significant constraint. The model cannot see your infographic. It can only read what you told it the infographic shows. This is also why W3C WAI's image guidance treats alternative text as content, not decoration: the same discipline that helps screen reader users also preserves meaning for extraction systems.
Tables depend on their markup. A properly marked-up <table> with <thead>, <th>, and <tbody> elements can be converted into a structured text representation that preserves the relationship between headers and data cells. A table built from <div> elements with CSS Grid, visually identical, becomes a flat sequence of text with no indicated structure. The model cannot reconstruct the grid from the CSS.
JavaScript-dependent content may not exist. If your page loads its main content through JavaScript (client-side rendering, infinite scroll, tab panels that load on click), that content may not be present in the HTML that the retrieval system fetches. Server-side rendering and static generation solve this problem completely; client-side-only rendering creates a page that looks complete in a browser and may look empty to a retrieval system.
What gets misinterpreted
Worse than losing information is having it misunderstood.
Navigation as content. Without <nav> landmarks, a retrieval system may include your site's navigation links, breadcrumbs, or footer links in the extracted "content." The model then treats "Home | About | Products | Blog | Contact" as part of the article, which pollutes the context and can produce confused responses that blend your navigation labels with your actual argument.
Repeated boilerplate as emphasis. If your header, sidebar, and footer all contain your company tagline or a call-to-action, and those elements are not separated from the main content by landmark elements, the model sees the same phrase three or four times. Repetition in a text input functions as emphasis: the model weights repeated phrases more heavily, which can skew its understanding of what the page is actually about. I have watched a model summarize a page's primary topic as the tagline from its header rather than the thesis from its opening paragraph, because the tagline appeared in three different locations while the thesis appeared once.
Heading hierarchy as topic structure. The model treats heading levels as an outline. If your page uses <h3> elements for visual reasons (you liked the size) rather than structural ones, the model interprets those as subsections of whatever <h2> precedes them. A page with random heading levels produces a nonsensical outline, and the model's understanding of the page's topic structure inherits that nonsense.
Missing definitions as assumed knowledge. When you use a technical term without defining it, a human reader might infer the meaning from context or look it up. A model processing your page for retrieval stores the term as-is. If another user later asks the model about that term, and your page used it without definition, the model may retrieve your page as a source while being unable to explain the concept clearly, reducing the quality of the citation. Entity clarity is not just good writing practice; it is a structural requirement for pages that want to serve as authoritative sources in retrieval systems.
What you can test right now
The gap between what you see and what a model sees is testable. You do not need to guess.
Quick audit: is your page machine-readable?
alt="") so they are correctly skipped.
Tools: axe DevTools, Accessibility Insights
Figure 3. Five tests you can run in under ten minutes to assess how well your page communicates with machines.
View the rendered text representation. Copy your page's URL into a tool that extracts readable text from HTML (Readability-based extractors, Mozilla's reader view, or browser developer tools' accessibility tree inspector). What you see is a rough approximation of what a retrieval system extracts. If the result is missing content, includes boilerplate, or has a broken structure, the model will encounter the same problems.
Read your page with a screen reader. VoiceOver (Cmd+F5 on macOS) or NVDA (Windows) will read your page the way the accessibility tree presents it. If the screen reader experience is confusing, fragmented, or missing content, the model's experience is likely similar. This is not a coincidence: both systems depend on the same structural signals. It is, in my experience, the single most illuminating test you can run on a page, because it forces you to hear your content the way a machine processes it rather than see it the way a designer intended it.
Check your heading outline. Browser extensions and developer tools can display the heading hierarchy of any page. The outline should read like a table of contents for the article. If it includes headings from navigation, sidebar widgets, or footer sections, those will leak into the model's understanding of the page's topic structure.
Disable CSS and JavaScript. The page with CSS and JavaScript disabled shows you something close to what the parser sees. If the page is unreadable, unstructured, or empty without styling, it will be unreadable, unstructured, or empty to the model.
Check your alt text. Every content image should have alt text that describes what the image communicates (not just what it depicts). Decorative images should have empty alt attributes (alt="") so they are correctly skipped. Missing alt attributes cause unpredictable behavior: some parsers skip the image, some include the filename, some include nothing.
The asymmetry you are building for
The core insight is an asymmetry: your users experience the rendered page, while an increasing share of your audience (answer engines, AI assistants, retrieval systems, and the humans who use them) experiences the parsed page. These are different things, and they can diverge dramatically.
A page can be visually beautiful and structurally illegible. It can have excellent typography, thoughtful layout, and a polished reading experience while being opaque to every system that reads the HTML rather than rendering it. The reverse is also true, though less common: a plain, unstyled page with strong semantic structure is more legible to a model than a visually sophisticated page built from non-semantic markup.
The practical resolution is not to choose one audience over the other. Semantic HTML serves both, because the structural clarity that makes a page accessible to screen readers and parseable by AI systems also makes the codebase more maintainable, the CSS more portable, and the page more resilient. The discipline of saying what things are, rather than only describing what they look like, is one of those rare cases where the principled approach and the pragmatic one converge. My four-year-old does something similar when she describes a room: she lists what matters (the dog, the TV, the snack on the table), ignoring everything decorative. The accessibility tree works the same way. So does an LLM's content extractor.
Open questions
We do not yet know how much semantic structure influences citation decisions in specific AI systems. Google has published guidance on structured data, though not on how semantic HTML specifically affects AI Overview inclusion. OpenAI has not published how ChatGPT's browsing feature weights structural signals versus content quality. Perplexity has disclosed some retrieval architecture details without describing its parsing heuristics. The experiments that would isolate the effect of semantic HTML on AI citation rates (serving semantically identical content in different HTML structures and measuring retrieval and citation differences) have not, to this writer's knowledge, been published.
What we can say with confidence is that semantic HTML makes it easier for these systems to correctly identify, extract, and structure your content. Whether "easier" translates to "more likely to cite" is a separate question, one that depends on content quality, authority signals, and system-specific ranking factors that operate above the parsing layer. The structural layer is necessary infrastructure, not a guarantee.
The experiment is worth running. We may run it ourselves.