AI Crawler
AI & Generative SearchAlso: AI Bot · LLM Crawler · AI Scraper
Quick definition
An AI crawler is an automated bot that reads your website so an AI system can either train on your content or cite it in a live answer. Named crawlers include GPTBot from OpenAI, ClaudeBot from Anthropic, PerplexityBot, and Google-Extended. You control access to each one through your robots.txt file.
How it varies across Australia
Most Australian sites still run the default robots.txt they shipped with, which means every AI crawler has full access whether the owner realises it or not. Awareness is rising faster in publishing and professional services than in retail, where the blocking decision is rarely on anyone's desk yet.
See data and tracking scores across Australian industries →The four crawlers you'll actually see
Gathers text mainly to train future GPT models.
Controlled via robots.txtFetches pages to train and to ground Claude's answers.
Controlled via robots.txtReads pages live so Perplexity can cite you in answers.
Block and you lose citationsSeparate control for Gemini training, split from Googlebot.
Blocking it won't hurt search rankWhat it actually means
An AI crawler is the same idea as Googlebot, an automated visitor that reads your pages, but with a different appetite. Where a search crawler indexes you so people can find your link, an AI crawler reads you so a model can either learn from your writing or quote you inside a generated answer.
That split matters because the two jobs use different bots, and you can allow or block each one separately. GPTBot and Google-Extended mostly gather text to train future models. ClaudeBot and PerplexityBot also fetch pages live to ground an answer they're writing right now. Some bots do both. The line is blurry and shifts with each vendor's policy.
Every one of these obeys robots.txt, the plain text file that tells crawlers where they may and may not go. So the blocking decision is genuinely yours. The catch is that it isn't one decision, it's several. Block GPTBot and your words won't train the next model. Block the crawlers that feed live answers and you vanish from AI search results entirely, the same way a robots.txt block would hide you from Google. Getting this right is now part of technical SEO, not a side issue.
Blocking the training bot to protect your content also blocks the door your future customers walk through.
How it shows up
AI crawlers show up first in your server logs, as user-agent strings like GPTBot, ClaudeBot, PerplexityBot and Google-Extended hitting your pages. If you've never looked, they're almost certainly already there.
The blocking decision shows up in your robots.txt file, where each crawler gets its own allow or disallow rule. It shows up again in the gap between what you intended and what you did, plenty of sites block GPTBot for training then wonder why they never appear in AI answers, not realising a separate live-fetch bot needed access for that.
And it shows up in AI search visibility, the growing share of buyers who ask an assistant instead of typing a query. If the crawler that grounds those answers can't read you, you're invisible in that whole channel.
The Australian context
Australia has no equivalent yet to the licensing deals large United States and United Kingdom publishers have struck with AI vendors, so most Australian businesses face the crawler decision with no commercial upside on the blocking side of the ledger. There's no cheque coming for saying no.
The Privacy Act reforms and the ACCC's ongoing scrutiny of digital platforms mean the regulatory picture will keep moving, but as of now blocking an AI crawler is a business and marketing choice, not a compliance one. For most Australian service businesses the practical question is simple. Do you want to be quotable when a customer asks an AI assistant for a recommendation in your category, or not.
Where people get this wrong
Related terms
Common questions
Should I block AI crawlers in my robots.txt?
For most businesses that want to be recommended, no. Blocking the live-answer crawlers removes you from AI search results your customers use. Publishers protecting licensable content have a stronger case. Decide per bot rather than applying one blanket block to everything.
How do I block a specific AI crawler?
Add a rule to your robots.txt naming the crawler's user agent, then disallow the paths you want protected. You can block GPTBot while allowing PerplexityBot, or the reverse. Each bot reads its own rule, so control is granular if you set it up that way.
Does blocking Google-Extended hurt my Google rankings?
No. Google-Extended controls whether your content trains Gemini and feeds some AI features. It's separate from Googlebot, which handles normal search indexing. You can block Google-Extended and keep full Google Search visibility. They're deliberately split for exactly this choice.
Will a robots.txt block actually stop AI from using my content?
Only for the crawlers that respect it, and only going forward. Robots.txt is an instruction well-behaved bots follow, not an enforced barrier. Content already scraped, or copied and republished elsewhere, sits outside its control. Treat it as a policy signal, not a lock.
Debrief
Get the next one
No spam. No fluff. Just the next article, straight to your inbox.
Keep exploring
About New Rebellion
New Rebellion is a marketing intelligence consultancy. We build tools, score Australian businesses on how their marketing actually performs, and publish Debrief every day. This dictionary is part of how we work in the open.
How we think →