Skip to main content

AI & LLMs

Which crawlers are recognised

The AI and search crawlers TrueStat recognises, and what each one is for.

The current list. It is extended as operators launch new crawlers; nothing is removed, so a name that appeared in your history stays readable.

By operator

OpenAI

User agentTypeVerifiable by IP
ChatGPT-UserAI assistantYes
OAI-SearchBotAI assistantYes
GPTBotAI trainingYes

Anthropic

User agentTypeVerifiable by IP
Claude-UserAI assistantYes
Claude-SearchBotAI assistantYes
ClaudeBotAI trainingYes
Claude-WebAI trainingYes
anthropic-aiAI trainingYes

Perplexity

User agentTypeVerifiable by IP
Perplexity-UserAI assistantNo
PerplexityBotAI trainingNo

Perplexity's published IP list has been stale since early 2025 and its traffic resolves to generic cloud addresses. So a Perplexity hit can be reported but never confirmed. See How reliable is it.

Google

User agentTypeVerifiable by IP
Google-CloudVertexBotAI trainingYes
GoogleOtherAI trainingYes
GooglebotSearch indexingYes
Storebot-GoogleSearch indexingYes
Google-InspectionToolSearch indexingYes

Googlebot and GoogleOther are separated deliberately. Googlebot indexes for search; GoogleOther collects for AI and research. Reporting them as one "Google" number would be the single most misleading thing this panel could do, since it would let ordinary search indexing masquerade as AI interest.

Microsoft Bing

User agentTypeVerifiable by IP
BingPreviewAI assistantYes
bingbot, msnbot, adidxbotSearch indexingYes

Apple

User agentTypeVerifiable by IP
ApplebotSearch indexingYes

Applebot serves Siri and Spotlight Suggestions.

Meta

User agentTypeVerifiable by IP
meta-externalagent, FacebookBot, meta-externalfetcherAI trainingNo

Meta publishes no IP ranges, so its traffic is permanently unverifiable.

ByteDance

User agentTypeVerifiable by IP
Bytespider, TikTokSpider, ByteSpiderAI trainingNo

Also publishes nothing.

Common Crawl

User agentTypeVerifiable by IP
CCBotAI trainingYes

Common Crawl trains no models itself, but its corpus is a training input for most of the models above — which is why it is on the AI panel rather than filed as a generic scraper.

Other named AI crawlers

Reported under Other: Amazonbot, cohere-ai, Diffbot, ImagesiftBot, omgili, YouBot, Timpibot, AI2Bot (AI training); DuckAssistBot, MistralAI-User, Kagi (AI assistant).

Other search engines

Reported under Other, as search indexing: YandexBot, Baiduspider, DuckDuckBot, Slurp (Yahoo), SeznamBot, Qwantify.

Social and link preview fetchers

facebookexternalhit, Twitterbot, LinkedInBot, Slackbot, Discordbot, WhatsApp, TelegramBot — reported under Other bots, not as AI. These fetch your page to build a link preview when someone shares it. Useful to see, but not AI traffic.

Anything else that identifies as a bot

A user agent containing bot, crawler, spider, scraper, HeadlessChrome, PhantomJS, curl/ or wget/ is recorded as Other bots.

Google-Extended and Applebot-Extended

Neither will ever appear on this panel, and their absence is not a gap in our coverage.

They are not crawlers. They are robots.txt control tokens. Google's own documentation states that Google-Extended "doesn't have a separate HTTP request user agent string" — it governs whether content Google has already crawled may be used to train Gemini. Applebot-Extended is the same for Apple's models.

Neither one makes a single request. There is nothing to count.

So if you have blocked Google-Extended in your robots.txt and are checking whether that worked, this panel cannot tell you — nothing would have appeared either way. What you would look at instead is GoogleOther, which is a real crawler that does make requests.

Any row claiming to be Google-Extended would be either a spoofer or a bug in our classifier, which is why the database refuses to store the value at all rather than trusting the classifier to keep getting it right.

When something is missed

A crawler we do not recognise is counted as a human.

That direction is deliberate. It understates the AI panel rather than inflating it, and for a feature sold on honesty, understating is the correct way to be wrong.

If you are seeing traffic you believe is a crawler and it is not showing here, send the user-agent string to support@truestat.io and it can be added.

Was this page helpful?

Last updated August 28, 2026