The short version
You do not need llms.txt. No major AI system has confirmed it reads the file, nobody has produced traffic data showing it changes citations, and a simple test called "cats.txt" quietly exposed how thin the evidence really is. If you run a small site and someone is telling you llms.txt is a GEO must-have, they are selling you a horoscope.
The good news: the work that genuinely gets your pages quoted by ChatGPT, Perplexity and Google's AI answers is the same boring on-page work that already helps you rank. Clear structure, direct answers, clean HTML. You almost certainly have most of it already.
Below is why the four usual arguments for llms.txt fall apart, what cats.txt proved, and where to spend your afternoon instead.
What llms.txt is supposed to do
llms.txt is a proposed plain-text file you drop at your site root that hands large language models a curated, Markdown-friendly map of your best content. The pitch is that AI crawlers will read it, prefer it, and cite you more often as a result.
It sounds reasonable. A robots.txt for the AI era, a tidy summary so the machines don't have to wade through your navigation and cookie banners. The problem is that "sounds reasonable" is not evidence, and after more than a year of enthusiastic adoption there is still no confirmation from OpenAI, Anthropic, Google or Perplexity that any of them fetch the file, let alone weight it.
The four arguments for llms.txt, and why each one folds
Every case made for llms.txt reduces to one of four claims, and none of them survives contact with actual server logs.
1. "The big AI companies support it"
They don't say they do. Support means a public statement or documented crawler behaviour. What we have instead is a community spec and a lot of hopeful blog posts. Absence of a denial is not a yes.
2. "My server logs show bots fetching llms.txt"
A fetch is not a benefit. Bots crawl all sorts of speculative paths, and a request for the file tells you nothing about whether the content inside it influenced a single answer. Correlating "I added the file" with "I got cited" ignores everything else that changed on your site that month.
3. "It can't hurt"
It can cost you time, and it can drift out of sync. A static file that duplicates your content becomes a second thing to maintain. Let it go stale and you are now feeding a curated but wrong summary of your own site.
4. "Everyone's doing it, so it must work"
Adoption is not evidence. Plenty of sites added meta keywords for a decade after Google stopped using them. Popularity measures marketing, not results.
What the cats.txt test actually proved
The cats.txt test showed that the "AI reads my curated file" belief is unfalsifiable the way it's usually argued. The idea, surfaced in a Search Engine Journal piece, was simple: if people credit llms.txt for citations without a control, you can invent an equally fictional file, "cats.txt", and build the same kind of just-so story around it. There is no mechanism, no confirmation, and no clean before/after that isolates the file from everything else.
The point isn't that llms.txt is technically impossible to be useful someday. It's that today's argument for it is astrology: a pattern read into noise, immune to being disproven because no one runs a proper test. Correlation, a friendly narrative, and a file that happens to exist.
If you want to believe in it, at least test it honestly. Add the file to half your pages' worth of content, leave the other half alone, and watch citations over 60–90 days. Nobody selling llms.txt has shown you that experiment, because when it's been attempted the signal isn't there.
What actually moves AI citations
AI systems cite the pages they can read, trust and lift a clean answer from, which is ordinary on-page quality, not a special file. LLM-powered search still leans on the crawlable web, your existing structured content and your reputation across other sites. That's where your effort pays off.
Here's the work that reliably helps, in rough order of payoff:
- Answer the question in the first two sentences. Open each section with a direct, self-contained answer, then explain. This is the exact chunk an AI model extracts and quotes.
- Use real heading structure. One h1, descriptive h2s phrased as the questions people ask. Machines and humans both navigate by them.
- Add Schema.org markup. FAQPage, Article, Product, Organization. It removes ambiguity about what your page is and who wrote it.
- Keep the HTML clean. Server-render your main content. If your answer only appears after JavaScript runs, assume some crawlers never see it.
- Earn mentions elsewhere. Being referenced and linked from other credible sites is still the strongest trust signal, and AI answers lean on it heavily.
- Fix the boring technical stuff. Fast responses, no broken links, an accurate sitemap.xml, and a robots.txt that doesn't accidentally block the crawlers you want.
None of this is new. It's the same discipline that earns featured snippets, and AI answers are drawing from the same well.
A five-point checklist worth more than any llms.txt
If you have one afternoon, spend it here instead of authoring a file no one has confirmed they read.
| Do this | Why it pays off |
|---|---|
| Rewrite the first line of your top 10 pages as a direct answer | Gives AI a clean, quotable snippet |
| Add FAQ schema to your key pages | Makes your Q&A content machine-readable |
| Check your content renders without JavaScript | Ensures crawlers actually see it |
| Confirm robots.txt isn't blocking AI or search bots by accident | You can't be cited if you're not crawled |
| Get one solid mention or link from a trusted site in your niche | Strongest signal AI answers rely on |
Hosting sits underneath all of this. If your pages are slow to respond or your server intermittently drops crawlers, none of the on-page work lands. At TPC Hosting we keep things EU-hosted and GDPR-friendly, and there are real engineers on support 24/7 if you need a second pair of eyes on a robots.txt rule or a rendering issue. That's a far better use of your time than tending a speculative text file.
So should you ever add llms.txt?
Add it only if you enjoy tidy experiments and you commit to keeping it accurate. It won't hurt a small site to have one, but treat it as a maybe-someday bet, not a priority, and never at the expense of the checklist above. If a "GEO expert" frames llms.txt as urgent or essential, that tells you more about their pitch than about how AI search works.
Spend your energy where the mechanism is known. Write clear answers, mark them up properly, host them somewhere fast and stable, and earn a bit of trust across the web. That gets you cited. A file you can't prove anyone reads does not.
FAQ
Does llms.txt improve my AI search visibility?
There's no evidence that it does. No major AI provider has confirmed it reads or weights the file, and nobody has published clean before/after data showing it changes citations. Your on-page content and site reputation are what actually get you quoted.
What is the cats.txt test?
It's a thought experiment showing the argument for llms.txt is unfalsifiable as usually made. By inventing an equally fictional 'cats.txt' file, you can construct the same correlation-based story that credits llms.txt for citations, which exposes that the reasoning relies on narrative rather than a controlled test.
If llms.txt can't hurt, why not just add it?
Because it's a maintenance cost with no confirmed payoff. A static file that duplicates your content can drift out of sync and end up feeding a wrong summary of your own site. If you add it, commit to keeping it accurate and don't treat it as a priority.
What genuinely gets my pages cited by AI answers?
Clear, directly-answered content that crawlers can read plus trust signals from other sites. Open sections with a self-contained answer, use proper headings and Schema.org markup, server-render your main content, and earn mentions from credible sites in your niche.
Could hosting affect whether AI systems cite my site?
Yes, indirectly but meaningfully. If your server is slow or intermittently drops crawler requests, your pages may not be indexed or read fully, so none of your on-page work lands. Fast, stable hosting is the foundation everything else sits on.

