Chase Konchak

OpenAI Says robots.txt Might Not Stop Its Bot

handwritten sign taped to glass shop door

OpenAI keeps documentation for its web crawlers, and there's a line in it I had to read twice. For the bot called ChatGPT-User, OpenAI says that because those visits start with a real person asking a question, robots.txt "rules may not apply." That's a major AI company saying, in writing, that it might ignore part of the file you use to keep bots out.

Search Engine Journal covered it Friday, and they paired it with numbers from TollBit's bot report for the first half of 2026: about 15% of the AI page fetchers on European sites reached pages those sites had marked off limits, and ChatGPT-User did it on more sites than any other bot in the study.

So what is robots.txt actually doing?

Picture a paper sign taped to the front door of your shop. It says the showroom is open and the back office is staff only. Anyone who can read it will follow it if they feel like following it, because there's no lock and no bouncer behind that sign.

That's basically what robots.txt is. It's just a text file sitting at yoursite.com/robots.txt that names the bots and tells them which pages to stay out of. Google follows it, Bing follows it, and the AI companies say their training crawlers follow it.

And honestly, I had this wrong for a long time. I thought something actually enforced the file, like the server reads it and turns bots away at the door. No. Nothing checks it. Your server hands the page to whoever asks for it, and the file is just a record of what you asked for. So when OpenAI says the rules may not apply to one of its bots, there's nothing on your end that makes them apply.

OpenAI runs three bots, and they do three different jobs

GPTBot is the one collecting pages to train models. Block it and your content stays out of the training data.

OAI-SearchBot is the one that decides whether your business shows up when ChatGPT answers somebody's question. OpenAI's docs say sites that opt out won't appear in search answers, so this is the bot tied to getting found.

ChatGPT-User is the errand runner. Someone asks ChatGPT what you charge, and ChatGPT goes and reads your pricing page right then. This is the one OpenAI says robots.txt may not cover.

Three different bots with three different outcomes, and nothing warns you when you block the wrong one.

The mistake that costs money

Here's how it usually goes. A business owner reads an article called something like "block AI bots from stealing your content," pastes a stack of rules into robots.txt, and closes the tab feeling like they handled it. A lot of those pasted lists include OAI-SearchBot.

Say a dentist did that back in March. A woman across town opens ChatGPT and types "dentist near me who takes Delta Dental and has Saturday hours." ChatGPT goes looking, but the dentist's site is sitting behind a rule she pasted and forgot about, so the practice two miles away gets named instead. She blocked the bot that brings people in and left open the one she was actually worried about.

The TollBit numbers kind of show how much guessing is going on out here. 26% of North American sites block Anthropic's Claude-User, but only 9% of European sites do. Same bot, same web, totally different instincts about what to do with it.

What to check this week

Open yoursite.com/robots.txt in a browser. It's a public file, it takes ten seconds, and you'll see exactly what you're telling bots right now.

If you want people to find you through ChatGPT, and for a plumber or a dentist or a coffee shop that's a yes, let OAI-SearchBot through. Blocking GPTBot is a separate decision, because plenty of owners don't want their words feeding a training run, and that's fair. It costs you nothing in visibility, so the two decisions don't have to match.

Then look at your server logs or your CDN dashboard, because that's where you see which bots actually showed up and what they took. The robots.txt file records what you asked for. The logs show what happened.

The part nobody has settled yet

OpenAI put in writing that it may skip part of your file, and nobody is enforcing the rest of it either. So who's supposed to? I don't have an answer, and as far as I can tell, nobody does yet.

Meanwhile the traffic keeps climbing. Cloudflare's CFO told investors on their Q2 earnings call that machine traffic could hit 1,000 times human traffic within five years, and that it's not because humans are leaving the internet, it's just how fast the machine side is growing. Cloudflare expected machine traffic to pass human traffic in 2027. It happened in May.

None of that changes what you should do this morning. Check the file, and fix it if OAI-SearchBot is blocked. It's the cheapest thing on your list this week, and the expensive version of this mistake takes about a minute to undo.

We check this on client sites during the monthly pass, next to the analytics and the keywords that didn't land, because small settings drift. Somebody edits a file for a good reason, the reason expires, and the file just stays where they left it.

CONTACT

/ DIRECT CONTACT

Let’s get your phone ringing.

Let’s get your phone ringing.