{"uid":"cap_EdKsdwS2u9UpO_SdGy4bM","slug":"x402-agentutility-webpage-scraper-b8ca3f73","name":"x402 AgentUtility Webpage Scraper","description":"Scrape any webpage. Pulls title, description, canonical URL, OpenGraph + Twitter card metadata, headings, and outbound links from a single URL. Server-side rendering. Body content rendered as text / raw HTML / clean markdown. Optional link extraction. Cheerio-based, no headless browser — fast and cheap, ideal for static pages and SSR sites. Alias of scrape-website. For JS-heavy SPAs that need a real browser, see website-screenshot.","url":"https://x402.agentutility.ai/scrape?utm_source=zero.xyz","method":"POST","headers":{},"bodySchema":{"type":"object","$schema":"https://json-schema.org/draft/2020-12/schema","required":["input"],"properties":{"input":{"type":"object","required":["type","method","bodyType","body"],"properties":{"body":{"required":["url"],"properties":{"url":{"type":"string","description":"Public URL to fetch and parse. Must include scheme (http/https). Follows redirects."},"format":{"enum":["text","html","markdown"],"type":"string","description":"Body output format. 'text' (default), 'html' (raw), or 'markdown' (clean — best for LLM ingestion)."},"user_agent":{"type":"string","description":"Custom User-Agent header. Defaults to a modern desktop Chrome UA."},"include_links":{"type":"boolean","description":"If true, also returns an array of all <a href> links on the page. Default false."}}},"type":{"type":"string","const":"http"},"method":{"enum":["POST"],"type":"string"},"bodyType":{"enum":["json","form-data","text"],"type":"string"}},"additionalProperties":false},"output":{"type":"object","required":["type"],"properties":{"type":{"type":"string"},"example":{"type":"object","properties":{"h1":{"type":"string"},"og":{"type":"object","properties":{}},"url":{"type":"string"},"lang":{"type":"string"},"text":{"type":"string"},"title":{"type":"string"},"format":{"type":"string"},"twitter":{"type":"object","properties":{}},"canonical":{"type":"null"},"final_url":{"type":"string"},"body_chars":{"type":"integer"},"description":{"type":"string"},"status_code":{"type":"integer"}}}}}}},"responseSchema":{"type":"json","example":{"h1":"Example Domain","og":{},"url":"https://example.com","lang":"en","text":"Example Domain\n\nThis domain is for use in documentation examples without needing permission. Avoid use in operations.\n\nLearn more","title":"Example Domain","format":"text","twitter":{},"canonical":null,"final_url":"https://example.com/","body_chars":128,"description":"","status_code":200}},"example":null,"exampleRequest":null,"tags":["x402"],"displayCostAmount":"0.04","displayCostAsset":"USDC","priceDynamic":false,"priceHint":null,"priceStatus":"priced","priceSource":"registry","requiresHandshake":false,"reviewCount":0,"rating":{"score":"0.00","successRate":"0.00","reviews":0,"stars":null,"state":"unrated"},"availabilityStatus":"unknown","priceObserved":null,"sessionDeposit":null,"pricing":{"kind":"static","summary":"$0.04/call","primary":{"kind":"static","protocol":"x402","network":"base","amountUsd":"0.04","per":"call","confidence":"exact"},"accepted":[{"kind":"static","protocol":"x402","network":"base","amountUsd":"0.04","per":"call","confidence":"exact"}]},"paymentMethods":[{"uid":"pm_xIiKj4iexWfcbUAgLjMsU","protocol":"x402","methodType":"crypto","chain":"base","mode":"charge","costAmount":"0.04","costPer":"request","priority":0,"asset":null,"unit":"request","depositMicros":null,"planRef":null}],"brandName":null,"brandSlug":null,"brandBaseUrl":null,"brandDocsUrl":null,"whatItDoes":"Scrapes a single webpage and returns its title, metadata, headings, body content (text/HTML/markdown), and outbound links using Cheerio-based server-side rendering.","exampleAgentPrompt":"Can you scrape https://techcrunch.com/2024/01/15/openai-news/ and give me the page title, description, main headings, and body text as clean markdown?","exampleUseCases":[{"title":"SEO metadata audit for a URL","prompt":"Scrape https://www.shopify.com/blog/ecommerce-seo and pull out the page title, meta description, canonical URL, and all OpenGraph tags so I can review their SEO setup."},{"title":"Research article content extraction","prompt":"Fetch the full body text as clean markdown from https://www.nature.com/articles/s41586-023-06004-9 and also grab all the outbound links on the page."},{"title":"Competitor landing page analysis","prompt":"Scrape https://www.notion.so/product and give me the headings, body text, and any outbound links — I want to understand how they structure their product page."}],"resultDescription":"Returns structured data including page title, meta description, canonical URL, OpenGraph and Twitter card metadata fields, all heading tags (H1–H6), body content in the requested format (plain text, raw HTML, or clean markdown), and optionally a list of outbound hyperlinks found on the page.","failureModes":["URL is unreachable or returns non-200 HTTP status — endpoint returns an error with the HTTP status code","Page is a JavaScript-heavy SPA that requires a real browser — content may be empty or minimal since no headless browser is used","Malformed or invalid URL input — returns validation error","Page blocks server-side scraping via robots.txt or IP blocking — may return empty or error response","Payment failure (x402) — request is rejected before scraping begins"],"whenToPreferThis":"Choose this endpoint for static pages, server-side rendered sites, and any URL where content is available in the initial HTML response. It is significantly faster and cheaper ($0.04 USDC) than headless-browser alternatives. Avoid it for JavaScript-heavy SPAs where content is rendered client-side — use a browser-based screenshot or rendering service instead.","instructions":null,"reviewSummary":null,"reviewSummaryHighlights":null,"reviewSummaryConcerns":null,"reviewSummaryGeneratedAt":null,"activationCount":0,"lastUsedAt":null,"lastSuccessfullyRanAt":null,"lastHealthCheckAt":"2026-10-02T02:33:37.732Z","isFirstParty":false,"canonicalSlug":"x402-agentutility-webpage-scraper-b8ca3f73"}