{"uid":"cap_UFa6EA74vKxeahX09ruku","slug":"ai-crawler-policy-checker-a60abd16","name":"AI Crawler Policy Checker","description":"AI-crawler policy of a website: reads robots.txt and reports, for 20 known AI/LLM crawlers (GPTBot, ChatGPT-User, OAI-SearchBot, ClaudeBot, anthropic-ai, CCBot, Google-Extended, Applebot-Extended, PerplexityBot, Bytespider, Amazonbot, cohere-ai, meta-externalagent and more), whether the root is allowed, blocked or unmentioned, plus an overall stance (open / selective / blocks-all-ai) and TDM/ai.txt hints. $0.01 per domain.","url":"https://intel.rallylive.ca/dev/ai-crawler-policy","method":"GET","headers":{},"bodySchema":{"type":"object","$schema":"https://json-schema.org/draft/2020-12/schema","required":["input"],"properties":{"input":{"type":"object","required":["type","method"],"properties":{"type":{"type":"string","const":"http"},"method":{"enum":["GET"],"type":"string"},"queryParams":{"type":"object","properties":{}}},"additionalProperties":false},"output":{"type":"object","required":["type"],"properties":{"type":{"type":"string"},"example":{"type":"object"}}}}},"responseSchema":null,"example":null,"exampleRequest":null,"tags":["x402"],"displayCostAmount":"0.01","displayCostAsset":"USDC","priceDynamic":false,"priceHint":null,"priceStatus":"priced","priceSource":"probe","requiresHandshake":false,"reviewCount":0,"rating":{"score":"0.00","successRate":"0.00","reviews":0,"stars":null,"state":"unrated"},"availabilityStatus":"unknown","priceObserved":null,"sessionDeposit":null,"pricing":{"kind":"static","summary":"$0.01/call","primary":{"kind":"static","protocol":"x402","network":"base","amountUsd":"0.01","per":"call","confidence":"exact"},"accepted":[{"kind":"static","protocol":"x402","network":"base","amountUsd":"0.01","per":"call","confidence":"exact"}]},"paymentMethods":[{"uid":"pm_ldGuQQVhzyUMQUiF_fjER","protocol":"x402","methodType":"crypto","chain":"base","mode":"charge","costAmount":"0.01","costPer":"request","priority":0,"asset":"0x833589fCD6eDb6E08f4c7C32D4f71b54bdA02913","unit":"request","depositMicros":null,"planRef":null}],"brandName":null,"brandSlug":null,"brandBaseUrl":null,"brandDocsUrl":null,"whatItDoes":"Reads a domain's robots.txt and reports whether 20+ known AI/LLM crawlers are allowed, blocked, or unmentioned, plus an overall AI access stance and TDM/ai.txt hints.","exampleAgentPrompt":"Can you check openai.com's AI crawler policy — specifically which AI bots like GPTBot, ClaudeBot, and PerplexityBot are allowed or blocked in its robots.txt, and what its overall AI access stance is?","exampleUseCases":[{"title":"Vetting a data source for AI training","prompt":"Before we scrape articles from theguardian.com for our LLM training dataset, can you check its AI crawler policy and tell me if bots like CCBot and anthropic-ai are allowed or blocked?"},{"title":"Competitive AI policy benchmarking","prompt":"I want to compare how openai.com, anthropic.com, and google.com treat AI crawlers — can you pull the AI crawler policy for each and tell me their overall stance: open, selective, or blocks-all-ai?"},{"title":"Publishing compliance check before launch","prompt":"We're about to launch our new site and want to make sure our robots.txt correctly blocks all AI crawlers — can you check the AI crawler policy on mybrand.com and confirm every major LLM bot like GPTBot, ClaudeBot, and Bytespider is blocked?"}],"resultDescription":"Returns per-crawler permission status (allowed, blocked, or unmentioned) for up to 20 known AI/LLM crawlers including GPTBot, ChatGPT-User, OAI-SearchBot, ClaudeBot, anthropic-ai, CCBot, Google-Extended, Applebot-Extended, PerplexityBot, Bytespider, Amazonbot, cohere-ai, meta-externalagent, and more. Also returns an overall AI access stance (open, selective, or blocks-all-ai) and any TDM or ai.txt hints found on the domain.","failureModes":["Domain does not have a robots.txt — crawlers reported as unmentioned with no stance","Domain is unreachable or returns non-200 HTTP status — endpoint may return error or empty result","Invalid or malformed domain input — request fails with validation error","robots.txt is present but uses non-standard syntax — some rules may be misclassified","Rate limiting on the target domain may prevent robots.txt fetch"],"whenToPreferThis":"Use this endpoint when you need a structured, per-crawler breakdown of a domain's AI access policy based on robots.txt, especially when you need to distinguish between specific crawlers (e.g. GPTBot vs ClaudeBot) rather than a generic crawl-permission check. Prefer this over manually parsing robots.txt when you need the overall AI stance classification (open/selective/blocks-all-ai) or TDM/ai.txt signal in one call.","instructions":null,"reviewSummary":null,"reviewSummaryHighlights":null,"reviewSummaryConcerns":null,"reviewSummaryGeneratedAt":null,"activationCount":0,"lastUsedAt":null,"lastSuccessfullyRanAt":null,"lastHealthCheckAt":"2026-09-14T12:52:02.228Z","isFirstParty":false}