# Answer engines are welcome. Search and assistant crawlers are named explicitly # so the policy is stated, not inferred from the wildcard. They share ONE group # with "*" on purpose: a crawler that finds a group naming it ignores the "*" # group entirely, so a separate "Allow: /" group would silently drop every # Disallow below for that crawler. # OpenAI GPTBot · OAI-SearchBot · ChatGPT-User # Anthropic ClaudeBot · Claude-SearchBot · Claude-User # Perplexity PerplexityBot · Perplexity-User # Google Googlebot · Google-Extended (Gemini grounding and training) # Microsoft Bingbot (also feeds Copilot and ChatGPT search) # Apple Applebot · Applebot-Extended # Common Crawl CCBot (the corpus most models learn from) User-agent: * User-agent: GPTBot User-agent: OAI-SearchBot User-agent: ChatGPT-User User-agent: ClaudeBot User-agent: Claude-SearchBot User-agent: Claude-User User-agent: PerplexityBot User-agent: Perplexity-User User-agent: Googlebot User-agent: Google-Extended User-agent: Bingbot User-agent: Applebot User-agent: Applebot-Extended User-agent: CCBot Allow: / # The connect door. An agent fetches this once to register the Vexa MCP server; # it is listed explicitly because the Disallow below would otherwise cover it if # it ever moved back under /api/, and because a stated Allow beats an inferred one. # The .txt form is what the prompt points at now — the extension tells the agent # it is fetching a file, not a page, before it has any header to read. The line # above already covers it by prefix; it is spelled out because an agent matching # literally should not have to reason about prefixes to open its own door. Allow: /connect/redeem Allow: /connect/redeem.txt # The short summary written for language models. Allow: /llms.txt Disallow: /api/ Disallow: /email-preview/ Disallow: /email-verification/ Disallow: /dashboard/ Disallow: /ga-test/ Disallow: /test-cover/ Sitemap: https://vexa.ai/sitemap.xml