Skip to main content
Answer Engine OptimizationBy Kevin O'Connell13 min readOctober 1, 2026

Do B2B SaaS Companies Publish an llms.txt? October 2026 Data

390 B2B SaaS sites, one test: does the site serve a valid llms.txt? About half do, but a quarter of those files lack the spec's full structure.

About half. Of the 374 B2B SaaS sites we could check on October 1, 2026, 192 serve a valid llms.txt at their root (51%), and 64% do once docs subdomains count. The files vary a lot, and publishing one says nothing about whether an AI system reads it.

  • Established ahead. 59.6% of established brands publish one, against 46.1% of Y Combinator companies
  • A low bar. One in four valid files lacks the spec's full shape, and 23 have no links
  • Docs hide a lot. 40 sites publish only on a docs subdomain
  • Linked to robots.txt. 73% of sites that name AI crawlers there publish one, against 48% of the rest (an association)

How many B2B SaaS sites publish a valid llms.txt?

About half. Of the 374 sites in our panel we could check, 192 (51.3%) serve a valid llms.txt at the root. Established brands lead Y Combinator companies, 59.6% to 46.1%, and counting docs subdomains lifts the total to 64.4%.

llms.txt among B2B SaaS sites, October 2026
390 sites checked on October 1, 2026. Rates are shares of the sites we could check.
51.3%
serve a valid llms.txt at the root
192 of 374 sites. 95% interval 46.3% to 56.4%
59.6%
of established brands
Against 46.1% of Y Combinator companies (87 of 146 and 105 of 228)
64.4%
when docs subdomains count too
232 of 360 sites. 40 publish only on a docs subdomain
16
of 390 sites left out
4 domains no longer resolve and 12 never gave our checker a clear answer. The rate stays between 49.7% and 52.8% either way
Valid llms.txt, with 95% intervals
The line on each bar is the 95% interval: the range the true rate for this panel plausibly sits in.
Established brands59.6% 87 of 146 (51.5% to 67.2%)
Y Combinator companies46.1% 105 of 228 (39.7% to 52.5%)
Both cohorts, site root only51.3% 192 of 374 (46.3% to 56.4%)
Both cohorts, root or docs subdomain64.4% 232 of 360 (59.4% to 69.2%)

"Valid" has one test, the one the llms.txt specification treats as required: the file opens with a title, a markdown H1. We counted a site only when the response was a text file that passed it, not an HTML page that happened to return a 200. Four panel domains no longer resolve and 12 sites never gave our checker a clear answer, mostly because of bot walls, so they sit outside the rates. Whichever way those 12 would fall, the rate stays between 49.7% and 52.8%.

The established cohort leads by 13.5 points (95% interval 3.2 to 23.5), a gap too large to put down to chance alone. That cohort is hand-picked from recognizable brands. The Y Combinator cohort is a filter on a public directory of growth-stage companies. Treat both as panels, not as a census of B2B SaaS.

Counting docs subdomains too

Forty sites publish an llms.txt only on a docs subdomain such as docs.company.com, and nowhere at the root. DigitalOcean, Docker, Snowflake and LaunchDarkly are four of them. Counting those files, 232 of 360 sites (64.4%) publish somewhere we looked. A study that probes only the site root misses all 40.

The specification allows a file at the root or at any subpath, so one more convention was worth a check: /docs/llms.txt on the main domain. None of the 147 sites without a root or docs-subdomain file had a valid file there, so it adds nothing in this panel. We added that check after seeing the first results, and other subpaths were not probed.

Counting only the site root misses about 13 points: 40 B2B SaaS sites publish an llms.txt on a docs subdomain and nowhere else.

Why published llms.txt adoption rates disagree

Different populations. Each study counts a different one. Ranked lists of the biggest domains read 3.5% to 8.3%, Ahrefs' own customers 28%, UK B2B tech companies about a quarter, and a hand-picked panel of well-known hosts 57%. None tests the same thing.

Published llms.txt adoption figures, side by side
Each row counts a different population on a different date with a different test. This is context, not a ranking.
StudyWho was countedWhenShare with an llms.txtHow a file counted
Independent top-million crawlTop 1,000,000 domainsAug 20263.53%Plain-text response required
HTTP Archive analysisTop 10,000 sites (7,504 crawled)Jun 20265.61%Valid, parseable file; soft-404 pages screened out
GEO/AEO PlaybooksTop 5,000 domainsJul 20267.5%Spec-valid file; 1,590 non-responding hosts reported separately
RankabilityTop 10,000 domainsSep 20268.3%llms.txt or llms-full.txt; sites it could not check stay in the base
SE RankingNearly 300,000 domainsPublished Nov 202510.13%Not stated
Ahrefs137,210 Ahrefs Web Analytics sitesMay 202628%Root returns 200 with Markdown; Ahrefs calls it an upper bound
Marketing Graham audit, reported by IT Brief752 UK tech companiesReported Oct 2026About a quarterA genuine, non-empty file; method not published
llmtxt.info tracker219 hand-picked well-known hosts (216 tested)Sep 202657.4% (SaaS sector: 68.1%, 47 hosts)200, plain text, valid title; unreachable hosts left out
This study390 B2B SaaS sites (374 checkable)Oct 202651.3% (46.3% to 56.4%); established 59.6%, Y Combinator 46.1%Text file that opens with a title; unreachable and bot-walled sites left out

Nobody in the table is wrong. Ahrefs calls its 28% an upper bound because its customers skew technical, and ranked lists of top domains include infrastructure hosts that never serve a web page. The row closest to ours in method is the llmtxt.info tracker, which uses the same test (a 200, plain text, a valid title) and leaves unreachable sites out of the base. It reads 57.4% on a hand-picked panel of well-known hosts.

Our 51.3% describes a panel of sites in this category, picked from a brand list and a startup directory rather than found by crawling the web. It does not describe the web. SE Ranking's study, published in November 2025, is the oldest row, so read it as a baseline rather than today's figure.

Valid is the minimum: what the files contain

A title is enough. A valid file only has to open with one. Of our 192, 144 (75%) have the spec's full shape of title, summary and linked sections, and 23 have no links at all. The median file is 9.7 KB.

What the 192 valid files contain
Valid means the file opens with a title. The full spec shape adds a summary, a section and a link.
Valid: opens with a title192 100%
At least one link169 88.0%
Full spec shape: title, summary, a section and a link144 75.0%
No links at all23 12.0%
Our panel next to the web at large
Share of filesOur 192 valid filesCommon Crawl, 584,107 files
Full spec shape75.0%49.9%
No links at all12.0%22.6%
Different populations and similar, not identical, definitions of the full shape. Common Crawl's July 2026 crawl covers files from across the web.

The specification asks for a title, a short summary in a quote block, and sections of links. Only the title is required, so a one-line file passes. Of the 192 valid files, 155 (80.7%) have the summary, 184 (95.8%) have at least one section, 169 (88.0%) have at least one link, and 55 (28.6%) use the optional section for pages a reader can skip. 144 have all of it.

The 23 files with no links (12.0%) include several short notes addressed to AI crawlers rather than maps of the site, and one valid file is a for-sale notice from a parked domain. At the other end, 17 files (8.9%) run past 100 KB and the largest is 1.52 MB. The median file has about 76 lines and 48 links.

Ten more sites publish something that looks like an llms.txt but fails the title test, such as a list of links with no title, or a title preceded by a plugin credit line. We left them out of the 51%. Counting them would make it 54%.

Next to Common Crawl's July 2026 analysis of 584,107 files from across the web, our B2B SaaS files come out more complete on both measures in the comparison above. The populations differ, and so do the definitions of the complete shape, so read the contrast as a direction, not a precise gap.

Publishing is not always a choice. Common Crawl found that two thirds of the files it analyzed are templated, mostly by plugins and site builders. We did not measure generators, and only five of our 192 valid files carry a generator signature, so we draw no conclusion for the panel.

A valid llms.txt only has to open with a title. One in four of the valid files in our panel stops short of the spec's full structure.

Yes. 57 of 78 sites with an explicit AI-crawler rule (73.1%) also publish a valid llms.txt, against 133 of 279 without one (47.7%). It is an association, not a cause, and Cloudflare use shows no difference.

Valid llms.txt by robots.txt stance, with 95% intervals
An association between two scans 16 days apart, not a cause.
Name AI crawlers in robots.txt73.1% 57 of 78 (62.3% to 81.7%)
Do not name AI crawlers47.7% 133 of 279 (41.9% to 53.5%)
Served through Cloudflare51.7% 93 of 180 (44.4% to 58.9%)
Not served through Cloudflare51.0% 99 of 194 (44.0% to 58.0%)
The robots.txt comparison uses the 357 sites with a readable robots.txt on September 15, 2026 and a clear llms.txt answer. The Cloudflare comparison uses all 374 checkable sites. Within cohorts, Y Combinator companies: 73.9% against 41.6%. Established brands: 71.9% against 56.6%.

We scanned the robots.txt of the same 390 sites 16 days earlier for our AI Bot Blocking Index. A site counts as naming AI crawlers when its robots.txt mentions at least one of the 17 we track, whether it lets that crawler in or keeps it out. The comparison uses the 357 sites with a readable robots.txt and a clear llms.txt answer.

The link is strongest among Y Combinator companies: 73.9% (34 of 46) of those that name AI crawlers publish a valid file, against 41.6% (69 of 166) of those that do not. Among established brands it points the same way, 71.9% against 56.6%, but that gap is within what chance could produce (95% interval from -4.1 to +30.8 points).

Blocking is not the opposite of publishing. Of the 16 sites that block at least one tracked crawler, 11 publish a valid llms.txt. Sixteen sites is too few for firm conclusions, so treat that as description only. For which crawlers are worth letting in, see which AI crawlers to allow. Cloudflare use shows no link at all: 51.7% of sites served through Cloudflare publish a valid file, against 51.0% of the rest.

For the robots.txt side of the story, see what that scan found about blocking.

Where the well-known adopters keep their llms.txt

Mostly at the root. Vercel, Zapier, GitBook, Mintlify, Yoast and Mastercard's developer portal serve one there. Anthropic's is on its developer docs, and Hugging Face serves them on documentation subpaths rather than at its root.

Where well-known adopters keep their llms.txt
Checked October 1, 2026 with a plain request. A check of named companies, not a sample.
CompanyWhere the file isA file at the site root?
Vercelvercel.com/llms.txtYes
Zapierzapier.com/llms.txtYes
GitBookgitbook.com/llms.txtYes
Mintlifymintlify.com/llms.txtYes
Yoastyoast.com/llms.txtYes
Mastercard Developersdeveloper.mastercard.com/llms.txtYes, on its developer portal host
Anthropicdocs.anthropic.com/llms.txt, which redirects to platform.claude.com/llms.txtNo. anthropic.com/llms.txt is a 404
Hugging Facehuggingface.co/docs/hub/llms.txt and huggingface.co/docs/transformers/llms.txtNo. huggingface.co/llms.txt is a 404

Anthropic's file moved with its documentation: docs.anthropic.com/llms.txt now redirects to platform.claude.com/llms.txt, while anthropic.com/llms.txt returns a 404. Hugging Face's root returns a 404 too, but its Hub and Transformers documentation each serve a file.

Lists of llms.txt adopters tend to name companies without saying where each file lives. That matters because a check of the site root alone would miss Anthropic's and Hugging Face's files.

What to do if your B2B SaaS site has no valid llms.txt

Start with a title. Publish one file at your root with your company name as the title, a one-line summary and links to your key pages, served as plain text. If your docs live on a subdomain, publish there too. Then fetch it yourself to confirm.

  1. Pick the host. Publish at the root. If your documentation lives on a subdomain, publish there too: that is where 40 of our publishers keep the only file they have.
  2. Open with a title. The first line is # Your Company, followed by a one-line summary in a quote block, then sections of links with a few words about each page.
  3. Serve plain text. Return a 200 with a text content type, not an HTML page, and do not redirect the path to your homepage. Seventeen panel sites answered with an HTML page instead of a file.
  4. Fetch it yourself. Request the URL the way a crawler would and confirm the first line is the title. A catch-all page that returns 200 for every path looks published and is not.
  5. Keep it current. List the pages a buyer or an AI assistant would need first, and update the file when those pages move.

A minimal file in that layout (our llms.txt explainer has longer starters for SaaS, docs and blog sites):

llms.txt minimal example
# Blaze CRM

> Blaze CRM is a CRM for B2B sales teams, with pipeline reporting built in.

## Product

- [Pipeline reporting](https://blazecrm.com/reporting): Forecasts by stage.
- [Pricing](https://blazecrm.com/pricing): Plans and what each includes.

## Docs

- [API reference](https://docs.blazecrm.com/api): Every endpoint.

## Optional

- [Changelog](https://blazecrm.com/changelog): Release notes by month.

The free llms.txt Generator builds a file in this layout from a short form: a title, a summary and sections of links. No signup.

Build your file with the free llms.txt Generator →

llms.txt is one small piece of answer engine optimization. The free AEO audit includes an llms.txt check among its 29, and the AEO platform scores your site for AI citation with page-level recommendations.

What this study cannot tell you

Published, not read. This study measures what sites publish, not what AI systems fetch or use, and the panel is hand-picked, not random. Treat the rates as a description of these 390 sites.

  • Use. Google's own guide says Google Search does not use llms.txt, and that publishing one neither helps nor harms visibility there. In its May 2026 study, Ahrefs found that 97% of about 38,000 valid files on its customers' sites received no requests at all. Nothing here shows that publishing a file changes your AI citations.
  • The panel. 153 hand-picked established brands and 237 Y Combinator companies, so the rates belong to these sites. Some panel domains no longer run their own site, and ten valid files are served from a different domain than the one on our list: eight from a site that forwards to another domain, one from a GitHub gist and one from a website builder's file host. Counting those ten as not published gives 48.7%.
  • Where we looked. The site root, plus docs. and /docs/ for sites without a root file. A file on any other subpath is missed.
  • Silent sites. Twelve sites never gave our checker a clear answer, including on a re-check 32 minutes later. We left them out instead of counting them as misses. Two of the twelve did answer a browser's user agent, one with a valid file and one with none, and counting those two leaves the rate at 51.3%.
  • Intent. Plugins and documentation platforms write files by default, so a published file does not prove a team chose to publish it.
  • One snapshot. Everything was checked on October 1, 2026, and files change.

To see whether anything fetches your own file, read your own server logs, or track AI bot activity over time.

How we measured

One request per site. On October 1, 2026 we fetched /llms.txt from 390 B2B SaaS sites, 153 established brands and 237 Y Combinator companies, and counted a site as publishing only when the response was a text file that opens with a title.

The panel is the one behind our AI Bot Blocking Index: hand-picked established B2B SaaS brands, plus active B2B SaaS companies from the Y Combinator directory (status active, public or acquired; team size 10 or more). Each request was a plain GET of https://domain/llms.txt under our own checker's user agent, with redirects followed and the body read up to 2 MB.

Every site ends in one of five classes. Valid is a text file whose first meaningful line is a title. Invalid is a non-HTML file that fails that test. Absent includes 404s and HTML pages returned for the path. Inconclusive covers bot walls and errors that persisted after one re-check at least 30 minutes later, and unreachable means the domain does not resolve. The last two are left out of the rates and never counted as misses.

We fixed the question, the panel, the test and the analyses before the first scan, then checked our own scanner by re-fetching every site it had classed as valid, absent or invalid with a second tool (curl). That check found one mistake, a bot challenge counted as an empty file, so we fixed it and re-ran the whole scan. A few checks were added after we saw first results and are labelled that way: the /docs/llms.txt check, the count that sets aside files served from another domain, and the split of the robots.txt link by cohort. Intervals are 95% confidence intervals, and group comparisons use Fisher's exact test.

Every site and class is downloadable under CC BY 4.0 as a CSV file or a JSON file, with a data dictionary inside the JSON. The robots.txt columns come from the Index scan of September 15, 2026. Cite it as: AI-Advisors llms.txt Adoption Study, B2B SaaS panel, October 1, 2026 (CC BY 4.0). Our own llms.txt passes the same test.

Frequently Asked Questions

#How many B2B SaaS companies have an llms.txt?

About half of the ones we checked. On October 1, 2026, 192 of 374 checkable B2B SaaS sites (51.3%, 95% interval 46.3% to 56.4%) served a valid llms.txt at their root: 59.6% of established brands and 46.1% of Y Combinator companies. It is a fixed panel of 390 sites, so it describes those sites, not B2B SaaS as a whole.

#Which companies use llms.txt?

In our October 1, 2026 check, Vercel, Zapier, GitBook, Mintlify, Yoast and Mastercard's developer portal serve a valid llms.txt at the root of the host we checked. Anthropic serves one on its developer docs rather than anthropic.com, and Hugging Face on its documentation subpaths. These companies publish a file; we did not test whether anything reads it.

#Where should an llms.txt go, the root or a docs subdomain?

The specification allows a file at the root or at any subpath, and only the title is required. Put one at the root, and publish on your docs subdomain too if your documentation lives there. In our panel, 40 sites published only on a docs subdomain, which lifts the rate from 51.3% to 64.4% once those files are counted.

#What makes an llms.txt valid?

Under the llmstxt.org specification the only required element is a title: a markdown H1 as the first line. A short summary in a quote block, sections of links and an optional section are recommended. In our panel 192 files met the title test, and 144 of them also had a summary, a section and a link.

#Does llms.txt actually work?

Our study cannot say, because it measures what is published, not what is read. Ahrefs found that 97% of about 38,000 valid files on its customers' sites received no requests at all in May 2026. Whether any AI system relies on the file for its answers is a separate question that this data does not test.

#Does Google use llms.txt?

Not according to Google. Its guide to generative AI features in Search says Google Search does not use llms.txt, and that publishing one neither helps nor harms your visibility there. It adds that a file is fine to publish for other services that use these files. The guide covers Google Search only, not other AI systems.

#Can a site block AI crawlers and still publish an llms.txt?

Yes. They are separate files: robots.txt sets crawl rules and llms.txt is a content map. Of the 16 sites in our panel that block at least one of the 17 AI crawlers we track, 11 also publish a valid llms.txt. Sixteen sites is too few for firm conclusions, so treat that as description only.

Kevin O'Connell
Kevin O'Connell
Founder & AEO Consultant, AI-Advisors.ai

20-year B2B SaaS marketer. 3x Head of Marketing. One company exit (Sapling HR acquired by Kallidus, 2021). Now building AI-Advisors.ai to give mid-market B2B teams the AI visibility tools enterprise brands get. Writing about Answer Engine Optimization, ChatGPT Ads, Microsoft Copilot SEO, and the 5 A's of AI Marketing framework.

✦

Start tracking your AI visibility today

Install the tracking snippet, run your first audit, and see how AI platforms treat your brand. Start your 14-day free trial.

Get Started Free

Keep Reading

Answer Engine Optimization
Why Google Preferred Sources Is Your Newest AI-Visibility Lever
7 min read
Answer Engine Optimization
How to Become a Google Preferred Source: A B2B Playbook
8 min read
Answer Engine Optimization
What Is Google Preferred Sources? The New AI Visibility Lever
9 min read