AI Crawler

AI & Generative Search

Also: AI Bot · LLM Crawler · AI Scraper

What it isBots that read your site to train or answer AI
Named onesGPTBot, ClaudeBot, PerplexityBot, Google-Extended
The decisionBlock in robots.txt or let them in
Watch forBlocking training bots kills AI visibility too

Quick definition

An AI crawler is an automated bot that reads your website so an AI system can either train on your content or cite it in a live answer. Named crawlers include GPTBot from OpenAI, ClaudeBot from Anthropic, PerplexityBot, and Google-Extended. You control access to each one through your robots.txt file.

How it varies across Australia

Most Australian sites still run the default robots.txt they shipped with, which means every AI crawler has full access whether the owner realises it or not. Awareness is rising faster in publishing and professional services than in retail, where the blocking decision is rarely on anyone's desk yet.

See data and tracking scores across Australian industries

The four crawlers you'll actually see

GPTBot(OpenAI)

Gathers text mainly to train future GPT models.

Controlled via robots.txt
ClaudeBot(Anthropic)

Fetches pages to train and to ground Claude's answers.

Controlled via robots.txt
PerplexityBot(Perplexity)

Reads pages live so Perplexity can cite you in answers.

Block and you lose citations
Google-Extended(Google)

Separate control for Gemini training, split from Googlebot.

Blocking it won't hurt search rank

What it actually means

An AI crawler is the same idea as Googlebot, an automated visitor that reads your pages, but with a different appetite. Where a search crawler indexes you so people can find your link, an AI crawler reads you so a model can either learn from your writing or quote you inside a generated answer.

That split matters because the two jobs use different bots, and you can allow or block each one separately. GPTBot and Google-Extended mostly gather text to train future models. ClaudeBot and PerplexityBot also fetch pages live to ground an answer they're writing right now. Some bots do both. The line is blurry and shifts with each vendor's policy.

Every one of these obeys robots.txt, the plain text file that tells crawlers where they may and may not go. So the blocking decision is genuinely yours. The catch is that it isn't one decision, it's several. Block GPTBot and your words won't train the next model. Block the crawlers that feed live answers and you vanish from AI search results entirely, the same way a robots.txt block would hide you from Google. Getting this right is now part of technical SEO, not a side issue.

Blocking the training bot to protect your content also blocks the door your future customers walk through.

How it shows up

AI crawlers show up first in your server logs, as user-agent strings like GPTBot, ClaudeBot, PerplexityBot and Google-Extended hitting your pages. If you've never looked, they're almost certainly already there.

The blocking decision shows up in your robots.txt file, where each crawler gets its own allow or disallow rule. It shows up again in the gap between what you intended and what you did, plenty of sites block GPTBot for training then wonder why they never appear in AI answers, not realising a separate live-fetch bot needed access for that.

And it shows up in AI search visibility, the growing share of buyers who ask an assistant instead of typing a query. If the crawler that grounds those answers can't read you, you're invisible in that whole channel.

The Australian context

Australia has no equivalent yet to the licensing deals large United States and United Kingdom publishers have struck with AI vendors, so most Australian businesses face the crawler decision with no commercial upside on the blocking side of the ledger. There's no cheque coming for saying no.

The Privacy Act reforms and the ACCC's ongoing scrutiny of digital platforms mean the regulatory picture will keep moving, but as of now blocking an AI crawler is a business and marketing choice, not a compliance one. For most Australian service businesses the practical question is simple. Do you want to be quotable when a customer asks an AI assistant for a recommendation in your category, or not.

Where people get this wrong

Blocking all AI crawlers with one blanket robots.txt rule.Training crawlers and live-answer crawlers do different jobs. Blocking everything to protect content also removes you from AI search results your customers are already using.
Assuming a robots.txt block stops all AI use of your content.Robots.txt is a request, not a wall. Well-behaved bots obey it, but content already scraped, or reproduced by third parties, sits outside its reach.
Never checking which bots actually visit.Most site owners have never read their server logs or robots.txt, so their AI crawler policy is whatever the default shipped with, decided by no one.

Related terms

Common questions

Should I block AI crawlers in my robots.txt?

For most businesses that want to be recommended, no. Blocking the live-answer crawlers removes you from AI search results your customers use. Publishers protecting licensable content have a stronger case. Decide per bot rather than applying one blanket block to everything.

How do I block a specific AI crawler?

Add a rule to your robots.txt naming the crawler's user agent, then disallow the paths you want protected. You can block GPTBot while allowing PerplexityBot, or the reverse. Each bot reads its own rule, so control is granular if you set it up that way.

Does blocking Google-Extended hurt my Google rankings?

No. Google-Extended controls whether your content trains Gemini and feeds some AI features. It's separate from Googlebot, which handles normal search indexing. You can block Google-Extended and keep full Google Search visibility. They're deliberately split for exactly this choice.

Will a robots.txt block actually stop AI from using my content?

Only for the crawlers that respect it, and only going forward. Robots.txt is an instruction well-behaved bots follow, not an enforced barrier. Content already scraped, or copied and republished elsewhere, sits outside its control. Treat it as a policy signal, not a lock.

Debrief

Get the next one

No spam. No fluff. Just the next article, straight to your inbox.

Keep exploring

About New Rebellion

New Rebellion is a marketing intelligence consultancy. We build tools, score Australian businesses on how their marketing actually performs, and publish Debrief every day. This dictionary is part of how we work in the open.

How we think →