Compare site environments
Production and staging need opposite policies, which is precisely why they collide. Staging should be invisible. Production should be open to search crawlers with selective rules for everything else. The failure mode is a deployment that carries the staging file into the live environment, and it happens more than anyone admits. Search Engine Land describes an ecommerce team whose organic traffic dropped 90% within 24 hours after a routine Thursday push brought two lines of staging robots.txt to production.
The deeper problem is that robots.txt was never adequate protection for staging in the first place. John Mueller of Google recommends blocking access server-side instead, either through HTTP authentication or IP allowlisting, and he explains why robots.txt falls short: you need to remember to change it when moving from staging to production, which is another source of common problems, and URLs blocked this way can still get indexed without their content.
Put authentication in front of staging and the robots txt for AI crawlers question resolves itself. Then add a deployment check that fails the build if a production file contains a sitewide disallow. It's a few lines of shell script against a mistake that costs quarters of recovery.
Robots txt for AI crawlers
Now to the configuration itself. The approach that holds up is selective rather than categorical, and it starts from a position of openness that you narrow deliberately. Writing robots txt for AI crawlers works when each group answers a question you've actually asked.
Keep Googlebot and Bingbot fully accessible. That's the floor, and nothing about AI policy justifies touching it. Then handle the AI bot robots.txt groups by function. Training bots get whatever your organization decided about licensing, and that decision belongs to legal and leadership rather than to whoever is editing the file. Search-oriented agents stay open unless you've accepted losing citation eligibility in ChatGPT and Perplexity.
For sensitive or low-value paths, disallow them by directory rather than URL by URL. Consolidating rules at the directory level keeps you clear of the 500 KiB ceiling and makes the file readable six months from now. But remember what a disallow is worth here. If a path holds customer data or internal pricing, robots.txt is the wrong tool and listing the path publishes a map to it. Use authentication.
A workable production file for a site that permits AI search but not training reads like this:
1User-agent: *2Allow: /3 4User-agent: GPTBot5Disallow: /6 7User-agent: ClaudeBot8Disallow: /9 10User-agent: Google-Extended11Disallow: /12 13User-agent: OAI-SearchBot14Allow: /15 16User-agent: Claude-SearchBot17Allow: /18 19User-agent: PerplexityBot20Allow: /21 22Sitemap: https://example.com/sitemap.xml
Adjust the training decisions to match your policy. The structure is the point: named groups and no wildcard rule that quietly overrides your intent. Note that changes to search-related AI crawler access take time to register. OpenAI says robots.txt updates take about 24 hours to propagate through its systems, and Google caches the file for up to 24 hours as well.
Test before deployment
Syntax comes first. Run the file through a validator that applies the documented matching rules, then read it yourself, because a valid file can still be a wrong file. Google's open source robots.txt parser is the same library used in Search, and running it locally catches precedence errors in robots txt for AI crawlers that eyeballing the groups won't.
Then test at the URL level, once per named agent. Take your highest-value pages and check each one against every user-agent token in the file as well as the wildcard. The Search Console tester was retired in 2023, so the robots.txt report under Settings now shows fetch status and errors rather than letting you test a URL against a draft. Use the URL Inspection live test for Googlebot and a third-party validator for the AI bot robots.txt groups.
After deployment, verify the live file on every host and confirm the version Google fetched matches what you shipped. Then move to server logs. Log review is where AI bot robots.txt stops being theory, because you can see which agents actually arrived and whether a bot you disallowed kept coming anyway. Verify identity against published IP ranges rather than the user-agent string, since strings are trivially spoofed.
Keep watching for two weeks. Crawl stats and index coverage move on different clocks, and a blocked search bot shows up in coverage reports before it shows up in sessions.
Your rollback plan needs to be ready before you deploy:
-
Keep the previous file version in source control so restoring it is a single revert.
-
If a sitewide or high-value block goes live, restore first and diagnose afterward, then request a robots.txt recrawl in Search Console to shorten the cache window.
-
Set an alert on any production file containing Disallow: / under a wildcard group.
Access is not visibility
Here's the part that gets lost when the audit is finished and the file is clean. Permission only makes AI crawler access possible.
Allowing Googlebot doesn't guarantee indexing, and indexing doesn't guarantee ranking. Allowing OAI-SearchBot doesn't put you in ChatGPT answers, and being crawled doesn't make you citable. AI crawler access is a prerequisite you have to satisfy before any of the actual work counts, which is a different thing from being the reason a model quotes you.
The economics make this obvious. Cloudflare Radar data from the first quarter of 2026 shows ClaudeBot fetching roughly 23,951 pages per single referral, against a Googlebot baseline near 5:1. Enormous crawl volume, almost no traffic returned. Access was never the bottleneck for those sites.
What determines whether you get cited is content a model can extract cleanly and authority signals it can corroborate elsewhere. A permissive robots txt for AI crawlers removes an obstacle. Removing an obstacle is not the same as arriving.
Get an expert review
Snoika Foundation works on AI search visibility for B2B SaaS companies and digital products so they get found and cited across ChatGPT and Perplexity. That work starts with the technical layer of robots txt for AI crawlers, because a crawler that can't reach your pages can't recommend them either.
If you want a second set of eyes on your robots txt for AI crawlers before it ships, book a call with our team for a file audit and an implementation review that protects your organic search traffic.