Skip to main content

What ChatGPT, Claude and Gemini actually see of your site

A site can be perfectly readable in a browser and almost empty to an AI assistant’s crawler. The difference comes down to three things: what is served in the first HTML response, what is reachable, and what is extractable.

Christophe Bellec ·

Short answer

AI assistant robots fetch a page and read its text; most do not execute JavaScript, or execute it less well than Googlebot. Anything that only exists after client-side hydration can therefore be invisible to them. The deciding test is to fetch your page without running JavaScript and look at what remains: if the title, the introduction, the offer and the FAQ are not there, that is exactly what those robots see.

Three families of robot, three behaviours

Talking about “AI robots” in the singular is misleading: they do different jobs and they are not blocked the same way.

Training crawlers
They collect content to train models. Blocking them protects your content but has no effect on your visibility inside generated answers, since they play no direct part in them.
Search crawlers
They build the index the assistant queries when it needs current information. Blocking them removes you from the results those answers rest on.
On-demand fetchers
Triggered when a user asks a question or supplies a URL: they fetch the page right then. A citation often travels through them.

These three families are told apart by their user agent. A blanket `Disallow: /` in robots.txt blocks all of them indiscriminately, which is rarely what anyone wants.

JavaScript is the first breaking point

Googlebot can render a JavaScript page, with a delay and a limited budget. Most assistant robots do not, or do so far more narrowly. They read the HTML the server sends, and nothing more.

In practice, on a classic single-page application the initial HTML contains an empty `div` and a script. All the content (heading, offer, pricing, FAQ) only appears after execution. To those robots, the page is empty.

The fix is not to abandon React: it is to serve the essential content in the first response, through server rendering or static generation. The interactive part can stay on the client.

How to check in five minutes

  1. Fetch the page without running JavaScript: disable it in the browser, or use a plain `curl` on the URL from the command line.
  2. Look at the HTML you get: is the H1 there? The first paragraph? The list of offers? The FAQ?
  3. Check the links: are they `a` tags with an `href`, or clickable elements driven by JavaScript? Only the former are followed.
  4. Check the metadata: a unique `title`, a consistent `description`, an absolute canonical tag.
  5. Check response time: a robot that waits too long gives up.

If what you get at step 2 would not be enough for a human reader to understand what you sell, no answer engine will understand it either.

Readable is not enough: it has to be extractable

Once the text is reachable, the question becomes: can an engine isolate an exact statement from it without interpreting? This is where most corporate sites fail, not through a technical fault but through writing style.

“We support our clients with bespoke solutions tailored to their challenges” contains no extractable information. “Two- to ten-day architecture audit, deliverable: system map and prioritised action plan, remote or on site in the Paris area” contains six facts.

  • Answer the question in the first paragraph, then elaborate. Not the other way round.
  • Name things: durations, deliverables, scope, conditions, and what you do not do.
  • Write self-contained answers: understandable without the preceding paragraph, because that is how they will be read.
  • One intent per page. Two pages targeting the same question cannibalise each other.
  • A correct heading hierarchy: one H1, and H2s that pose real questions rather than slogans.

Structured data: useful, not magic

JSON-LD does not make your company appear in an answer. It removes ambiguity: who publishes, which entity stands behind the statement, how pages relate to each other, when it was published.

What matters more than the number of schemas is their consistency: stable identifiers, reused across pages, pointing at the same entity. Ten isolated schemas copied page by page are worth less than one properly connected graph.

And one absolute rule: markup must describe what is visible on the page. Markup that claims more than the content is ignored at best, penalised at worst.

The most common mistakes

  • Essential content rendered only on the client.
  • JavaScript navigation with no `a href` tags, leaving whole pages unreachable.
  • A `Disallow: /` inherited from a staging environment and never removed in production.
  • Relative canonical tags, or canonicals pointing at a different page than the page itself.
  • Language versions without reciprocal `hreflang`, competing instead of complementing each other.
  • Pages that all repeat the same generic description, so none of them answers a precise question.

What order to work in

Order matters, because each step conditions the next: reachable, then readable without JavaScript, then structured, then extractable, then tied to a clear entity. Working on the entity before fixing the rendering is pointless: nobody will read what you claim.

Frequently asked questions

  • Is a React or Next.js site penalised?

    Not in itself. The problem is essential content that only exists after JavaScript runs. Next.js with server rendering or static generation serves the content in the first response, which settles the question.

  • Should all AI robots be allowed in robots.txt?

    It is a trade-off. Blocking training crawlers protects your content without harming visibility; blocking search and fetch crawlers removes you from answers. These families are told apart by user agent, so the decision is made line by line.

  • Is JSON-LD enough to get cited?

    No. It clarifies entity identity and the relationships between pages, but an engine cites content, not markup. Visible, extractable content comes first; JSON-LD supports it.

  • How do I find out what engines understand about my site today?

    By fetching your pages without JavaScript, checking robots.txt, canonical and hreflang, and testing real questions about your subject and your brand. That is exactly what an SEO and GEO audit covers.

How I can help on this

Get in touch

More articles