Is Your Website Visible to AI?
AI visibility is whether AI assistants and answer engines can reach your site, read it, and name it in their answers. It is not the same as ranking on Google: a page can rank well and still be absent from AI answers. Visibility runs in three stages — reach (can their crawlers get to you), read (is your content in the HTML they fetch), and cite (does anything on the page earn being quoted). The first two you can check today.
Can an AI crawler get to your page at all, or does a rule or a firewall turn it away before it reads a word?
Once it arrives, is the content it needs actually in the HTML it fetches — or hidden behind scripts it does not run?
Given it can read you, does anything on the page earn being quoted by name in the answer a buyer receives?
What Does It Mean to Be Visible to AI?
Classic search visibility is about ranking a link in a list of results. AI visibility is about being named inside a written answer. The two overlap — clean, reachable, well-structured pages help both — but they are not the same outcome. An assistant does not hand back ten blue links; it composes a paragraph and names a few sources. Being one of those named sources is the new front row.
This is why a site can hold a strong ranking and still go uncited by every assistant a buyer asks. The mechanics that decide the two are related but distinct, and a page that has never been read by an answer engine cannot be named by one. The rest of this piece walks the three stages in order, because you cannot skip to the third. If an assistant is already leaving you out, a companion piece diagnoses why, layer by layer: Why Doesn’t ChatGPT Recommend My Business? Its measured counterpart is written out on how we measure.
Can AI Assistants and Answer Engines Reach Your Site?
Before anything can name you, it has to be allowed in. Assistants and answer engines send named crawlers to read the web, and your site decides which of them it welcomes. Most sites never make that decision on purpose — which is exactly where the trouble starts.
Two files govern the front door. Your robots file, at yourdomain.com/robots.txt, tells crawlers what they may read; a newer companion, described further down, speaks to assistants directly. A single overly broad line in the robots file can turn away every AI crawler at once — the accidental total block, and the most common reason a site that means to welcome AI is invisible to it. It is rarely a deliberate choice; it is one rule written more broadly than intended. A server firewall can do the same thing lower down, refusing a reader before the robots file is ever consulted.
OpenAI runs two named readers worth knowing: GPTBot, which collects pages to help train its models, and ChatGPT-User, which fetches a page in the moment someone asks ChatGPT about it. The wider field is below. Each crawler in the table is one the free scan checks your robots posture against, read straight from the scanner so the two stay in step.
| Crawler | Operated by | What it reads for |
|---|---|---|
| GPTBot | OpenAI | Collects public web pages to help train and improve OpenAI’s models. |
| ClaudeBot | Anthropic | Gathers public pages to help train and improve Anthropic’s assistant. |
| PerplexityBot | Perplexity | Indexes public pages so they can be surfaced and cited in Perplexity’s answers. |
| CCBot | An open web-archive project | Gathers public pages into an openly published archive that many AI training sets are built from. |
| Google-Extended | A robots setting that governs whether your content helps train and ground Google’s generative answers — separate from ordinary Search indexing. | |
| Applebot | Apple | Powers Siri and Spotlight results; a companion setting governs whether your content helps train Apple’s models. |
| Amazonbot | Amazon | Collects public pages for Amazon’s services and its assistant. |
| Bytespider | ByteDance | Collects public pages for ByteDance’s AI products. |
Operator and purpose are public facts from each operator’s own crawler documentation. The checked-crawler set is read from the scan itself.
Can AI Crawlers Read What Your Page Actually Says?
Reaching a page is not the same as reading it. Many crawlers fetch your raw HTML and stop there — they do not run the scripts a browser runs. If the words that matter appear only after JavaScript builds the page, a non-rendering crawler sees an empty frame where your content should be.
- Your key content is present in the raw HTML, not assembled by scripts.
- One clear main heading states, in a line, what the page is about.
- Real, visible dates and a clean, sectioned structure a machine can follow.
- Descriptive link text, so the paths through your site say where they lead.
- The main content painted in only after a script runs.
- No main heading, or several competing for the role.
- Undated pages and a structure held together by styling alone.
- Vague links a machine cannot map to a destination.
You can see your own page the way a non-rendering crawler does in one step: open it and view its source. If the sentence you most want an assistant to quote is not in that raw markup, the crawler that stopped there never had it.
What Makes an AI Answer Name Your Site?
Reachable and readable get you considered. Being named is earned by what a page actually offers the answer. This is the stage no access check can measure, and the one that decides whether you appear in the reply a buyer reads.
- A direct answer to a real question, stated plainly near the top.
- Extractable sentences that stand on their own out of context.
- Facts marked up as structured data an engine can lift without guessing.
- Signals of a real author, real expertise, and real citations behind the claim.
- Being named elsewhere — the third-party mentions on-page work cannot replace.
- A page that circles a topic without ever answering the question.
- Claims with nothing extractable an engine can quote cleanly.
- Facts buried in prose with no machine-readable form.
- Confident assertions with no author, evidence, or corroboration.
Which of these your buyers actually reward is a market-specific question, and the honest answer is that it varies by category and city. The full Rank Codex answers it directly: it puts real buyer questions to the answer engines and records who gets named — the measured form of this stage.
What Is the AI Counterpart to the Robots File?
A newer convention, llms.txt, is a small plain-text file a site publishes to describe itself to AI assistants — what it is, which pages matter, and the terms on which its content may be used. Where the robots file tells crawlers what they may read, llms.txt speaks to the assistant that reads them.
It is early. It is a convention, not a settled standard, and support varies from one assistant to the next. But it costs little to publish and it states your intent in a form built for the readers now shaping how buyers find you. This site publishes one, as a working example, at rankassurance.com/llms.txt; how our own reader behaves is documented at our crawler.
How Can You Check Your AI Visibility Today?
You can test the first two stages by hand, right now, with nothing but a browser and an assistant.
Ask an assistant a question your buyer would ask — the best of what you do, in your city — and read who it names. If you are not among them, that is your visibility as a buyer meets it.
Open yourdomain.com/robots.txt and see whether it welcomes the crawlers above or turns them away. One broad line is all it takes to close the door.
Open your key page and view its source. If the words that matter are not in the raw HTML, a non-rendering reader never had them.
The free scan runs the mechanical half of that self-test for you. One of its phases reads your homepage the way a non-rendering crawler does and checks your robots posture for each of the crawlers above — reporting what it finds, every finding carrying how it is known, and an honest absence wherever a thing cannot be measured. It is the instrument behind the manual check.
What Being Reachable Does and Does Not Prove
Being reachable is the price of entry, not the prize. It means a crawler can read you; it does not mean an answer engine will name you.
Everything you can verify mechanically lives in the first two stages, and this is why. A check that reports a clean bill on access has told you the door is open. It has not told you whether anyone walks through it. Naming is earned by what the third stage rewards, it varies by market, and it is never guaranteed. Anyone who promises otherwise is selling the one thing no one can honestly assure. We measure what can be measured, we label what is estimated, and we refuse to invent the rest — the same posture behind every RankAssurance report.
Notes on Sources
This is a technical explainer. Here is exactly what it rests on.
- This piece cites no statistics. It is a technical explainer, and inventing a figure here would break the one rule that matters most to us — we report only what can be measured.
- The crawler names and what each reads for are public facts, taken from each operator’s own published crawler documentation. Support and behaviour change over time; a named, dated page is how we keep that honest.
- The set of crawlers the free scan checks your robots posture against is read directly from the scanner, so the table on this page and the product itself can never drift apart.
Published 3 September 2026. The crawler landscape changes; this page is versioned by its date.
See Whether AI Can Read Your Site
A free reading of your homepage the way a crawler sees it, your robots posture for the named AI crawlers, and an honest absence wherever something cannot be measured.
