{"uid":"cap_YiEDopu8MGiEztSN2V3US","slug":"crawl-preflight-check-e7520d11","name":"Crawl Preflight Check","description":"Crawl pre-flight for a domain: will you be blocked, is content behind a paywall, and how much is actually there. Returns per-bot robots.txt verdicts for 16 AI crawlers, crawlable page count from sitemaps, detected CDN and paid-access signals (x402, TollBit), plus a fetch/skip recommendation. Call before spending requests on an unknown domain.","url":"https://crawl-preflight.postnov01.workers.dev/crawl-check/:domain","method":"GET","headers":{},"bodySchema":null,"responseSchema":null,"example":null,"exampleRequest":null,"tags":["x402"],"displayCostAmount":"0.02","displayCostAsset":"USDC","priceDynamic":false,"priceHint":null,"priceStatus":"priced","priceSource":"registry","requiresHandshake":false,"reviewCount":0,"rating":{"score":"0.00","successRate":"0.00","reviews":0,"stars":null,"state":"unrated"},"availabilityStatus":"unknown","priceObserved":null,"sessionDeposit":null,"pricing":{"kind":"static","summary":"$0.02/call","primary":{"kind":"static","protocol":"x402","network":"base","amountUsd":"0.02","per":"call","confidence":"exact"},"accepted":[{"kind":"static","protocol":"x402","network":"base","amountUsd":"0.02","per":"call","confidence":"exact"}]},"paymentMethods":[{"uid":"pm_3ArTbDGzf27-ba9lw4H7c","protocol":"x402","methodType":"crypto","chain":"base","mode":"charge","costAmount":"0.02","costPer":"request","priority":0,"asset":null,"unit":"request","depositMicros":null,"planRef":null}],"brandName":null,"brandSlug":null,"brandBaseUrl":null,"brandDocsUrl":null,"whatItDoes":"Checks a domain before crawling to determine if AI bots are blocked by robots.txt, if content is paywalled, how many pages are crawlable, and whether to proceed or skip.","exampleAgentPrompt":"Before you start crawling techcrunch.com, run a preflight check to see if AI bots are allowed, whether there's a paywall, and how many pages are actually crawlable.","exampleUseCases":[{"title":"Pre-crawl access check for AI agent","prompt":"Before scraping any content from wsj.com, can you check if that domain allows AI crawlers, whether it's behind a paywall, and roughly how many pages we could actually access?"},{"title":"Vetting unknown domain before bulk crawl","prompt":"I want to crawl a list of domains starting with arxiv.org — first do a preflight check to see if it blocks bots like GPTBot, if there's any paid access barrier, and give me a fetch or skip recommendation."},{"title":"Detecting paywall signals before indexing","prompt":"Check substack.com before I add it to my indexing pipeline — I need to know if content is behind a paywall like TollBit or x402, how many sitemap pages there are, and what CDN they use."}],"resultDescription":"Returns per-bot robots.txt verdicts for 16 AI crawlers (e.g. GPTBot, CCBot), a crawlable page count derived from sitemaps, detected CDN provider, paid-access signals (x402, TollBit), and a clear fetch or skip recommendation.","failureModes":["Domain does not exist or is unreachable — may return error or empty result","robots.txt absent — verdicts default to allowed","Sitemap missing or malformed — crawlable page count may be zero or unreliable","Rate limiting on the target domain during preflight — partial data returned","Invalid domain format — request rejected"],"whenToPreferThis":"Use this endpoint when your agent is about to crawl an unknown or untrusted domain and needs to avoid wasting requests on blocked, paywalled, or content-sparse sites. It is especially valuable before bulk crawls, when building web indexing pipelines, or when you need per-bot robots.txt verdicts for specific AI crawlers rather than a generic accessibility check.","instructions":null,"reviewSummary":null,"reviewSummaryHighlights":null,"reviewSummaryConcerns":null,"reviewSummaryGeneratedAt":null,"activationCount":0,"lastUsedAt":null,"lastSuccessfullyRanAt":null,"lastHealthCheckAt":"2026-09-14T12:37:07.958Z","isFirstParty":false}