OpenAI's ChatGPT-User bot—the one that fetches pages when a user asks a question—may not be bound by your robots.txt rules. That's according to OpenAI's own documentation, which says that because a user initiated the request, the standard crawler rules may not apply. In the first half of 2026, TollBit's State of the Bots report found that ChatGPT-User accessed disallowed pages on nearly half of European sites that explicitly blocked it. This isn't just a technicality; it's a strategic shift in how AI companies treat publisher content.
For your business, this means your content can be pulled into ChatGPT answers even if you've told crawlers to stay out. But there's a nuance: OpenAI's search bot, OAI-SearchBot, which decides if your site appears in ChatGPT search results, does respect robots.txt. So blocking both bots to stop AI fetching also cuts you out of ChatGPT's search visibility. That's a trade-off you need to make consciously.
What the Data Shows: The Scale of the Bypass
TollBit's report, which tracks AI crawler behavior across European and North American sites, found that about 15% of AI page-fetchers reached URLs marked as disallowed. The worst offenders were ChatGPT-User, Bytespider (Baidu's bot), and Youbot (You.com's bot). Each accessed disallowed pages on nearly half of the European sites that had explicitly listed them. ChatGPT-User reached the most sites overall.
Interestingly, not all AI bots are equally blocked. Only 9% of European websites disallow Claude-User (Anthropic's bot), compared to 26% in North America. Perplexity-User sits at 13% in Europe versus 26% in North America. Most newer agents have single-digit disallow rates in Europe, but ChatGPT-User is the exception—it's blocked more often, yet still bypasses those blocks.
Why OpenAI's Stance Matters
OpenAI's documentation explicitly states that ChatGPT-User visits a page when a user asks a question, and because that action is user-initiated, robots.txt rules may not apply. Perplexity says its user bot generally ignores robots.txt for the same reason. Anthropic, however, says all three of its bots respect robots.txt. TollBit treats any request to a disallowed URL as a bypass, regardless of what the operator claims.
This creates a compliance gray zone. If you block ChatGPT-User, you're making a request that OpenAI says may not be honored. That's a significant shift from the traditional understanding that robots.txt is a binding instruction.
What This Means for Your Business
If you're a publisher or content owner, you need to decide: do you want your content to appear in ChatGPT's search results? If yes, you must allow OAI-SearchBot. But that also means ChatGPT-User can fetch your pages when users ask questions—even if you block it. If you block both, you lose search visibility but gain some control over fetching, though that control is uncertain.
For most small and mid-sized businesses, the practical impact is limited. If you're not heavily reliant on organic search traffic from AI assistants, you can ignore this. But if you're in a content-heavy industry—publishing, news, research—this is a strategic issue. You need to weigh the benefits of being cited in AI answers against the risk of your content being used without your consent.
One thing to watch: Cloudflare is updating its crawler controls. Starting September 15, new domains added to Cloudflare will have Training and Agent crawlers blocked by default on pages with ads, while Search crawlers remain allowed. This moves the decision to the network layer, meaning compliance is enforced by Cloudflare, not the crawler itself. This could become a standard for other platforms.
Strategic Responses for Content Owners
First, audit your robots.txt. Know exactly which bots you're blocking and what you're allowing. If you want to appear in ChatGPT search, allow OAI-SearchBot. If you want to minimize fetching, block ChatGPT-User—but understand it may not be honored.
Second, consider using Cloudflare's new controls. If you're on Cloudflare, you can enforce blocking at the network level, which is more reliable than relying on crawler compliance. This gives you a stronger position.
Third, monitor your server logs. Robots.txt only shows what you asked for; logs show what actually happened. If you see ChatGPT-User accessing disallowed pages, you have evidence to push back or negotiate.
Finally, think about your content strategy. If AI assistants are driving traffic to your site, that's valuable. If they're just extracting answers without sending clicks, you may want to restrict access. There's no one-size-fits-all answer.
The Competitive Landscape
Anthropic's strict compliance with robots.txt positions it as a publisher-friendly option. If you're choosing which AI tools to support, this could be a differentiator. Perplexity's admission that it ignores robots.txt for user-initiated requests may drive more sites to block it, potentially limiting its data sources.
OpenAI's stance is risky. It could lead to publisher backlash, legal challenges, and reduced access to high-quality content. But it also reflects a broader trend: AI assistants are becoming the new search interface, and they need real-time access to the web. The old rules may not fit.
Bottom Line
For most businesses, this is a wait-and-see situation. The industry is moving toward clearer standards, possibly through regulation or industry agreements. Cloudflare's move is a step in that direction. In the meantime, make informed choices about your robots.txt and monitor your logs. The power is shifting to content owners who control access to their data—use it wisely.
FAQ
Yes, according to OpenAI's documentation, because the request is user-initiated, robots.txt may not apply. This means your disallow rule for ChatGPT-User might not be honored.
It depends. Blocking ChatGPT-User may not be effective, and if you also block OAI-SearchBot, you lose visibility in ChatGPT search. Weigh the benefits of being cited versus the risk of unauthorized fetching.
ChatGPT-User fetches pages when a user asks a question, while OAI-SearchBot decides if your site appears in ChatGPT search results. OAI-SearchBot respects robots.txt; ChatGPT-User may not.


