CRAWLER REGISTRY

AI Crawler Directory: User-Agents & IP Ranges

A canonical reference of official web crawlers operated by OpenAI, Anthropic, Google, Perplexity, and Apple, including purpose, user-agent tokens, and documentation links.

User-Agent Vendor Purpose Description Reference
OAI-SearchBot OpenAI Real-Time Search & Product Feeds Powers web search in ChatGPT and crawls product images across CDN hosts for advertiser product feeds. Advertiser Docs ↗
GPTBot OpenAI Model Training Collects public data for future foundational model training. Docs ↗
Claude-SearchBot Anthropic Real-Time Search Retrieves web content for Claude live search answers. Docs ↗
ClaudeBot Anthropic Model Training Crawls public text for Claude training datasets. Docs ↗
PerplexityBot Perplexity Real-Time Search Searches and extracts source citations for Perplexity answers. Docs ↗
Google-Extended Google Training Governance Token used in robots.txt to control Gemini/Vertex training data collection. Docs ↗
AI SEARCH VISIBILITY → RECOMMENDATION OPPORTUNITY → CUSTOMER

Your customer asks AI ‘who should I choose?’ Is your website in the consideration set?

HTML&HTML prepares your website for visibility, citation eligibility and recommendation opportunity across AI search experiences. It shows measurable website-side blockers that can prevent discovery, understanding and source consideration.

01

BE DISCOVERABLE BY AI

robots.txt, sitemaps, canonicals, indexability and AI crawler access form the discovery foundation.

02

BE UNDERSTANDABLE

GEO, AEO, LLMO, entity graphs, schema and answer extractability reduce machine ambiguity.

03

BE SOURCE-READY

RAG/retrieval, original information, E-E-A-T, freshness and evidence support source eligibility.

04

TURN OPPORTUNITY INTO DEMAND

AAO, accessible journeys, intact links, measurable referrals and clear CTAs connect AI discovery to commercial action.

Recommendations, rankings, citations, traffic, customers and revenue are not guaranteed. HTML&HTML measures website-side technical and content blockers; it does not claim control over external AI systems.