{"uid":"cap_Iztjr3-W7bLhX4JUgYkqu","slug":"web-article-pdf-to-markdown-converter-fe561683","name":"Web Article & PDF to Markdown Converter","description":"Fetch a public article or PDF and return clean Markdown plus title, byline, siteName, excerpt and wordCount. HTML is extracted with Firefox reader-mode rules; PDFs return their text layer. Requires ?url=<public http(s) URL>. Errors: 400 missing_url|bad_url, 403 blocked_private (private and internal hosts refused, every redirect hop re-checked), 413 too_large above 8 MB, 422 no_text_layer for scanned PDFs or not_extractable for app shells, 504 fetch_timeout. To FIND urls, use /search.","url":"https://x402.donnyautomation.com/markdown?utm_source=zero.xyz","method":"GET","headers":{},"bodySchema":null,"responseSchema":{"type":"json","example":{"ts":"2026-08-01T00:00:00.000Z","url":"https://en.wikipedia.org/wiki/Markdown","title":"Markdown","byline":null,"format":"article","excerpt":"Markdown is a lightweight markup language…","finalUrl":"https://en.wikipedia.org/wiki/Markdown","markdown":"# Markdown\n\nMarkdown is a lightweight markup language…","siteName":"Wikipedia","truncated":false,"wordCount":3204}},"example":null,"exampleRequest":null,"tags":["x402"],"displayCostAmount":"0.01","displayCostAsset":"USDC","priceDynamic":false,"priceHint":null,"priceStatus":"priced","priceSource":"probe","requiresHandshake":false,"reviewCount":0,"rating":{"score":"0.00","successRate":"0.00","reviews":0,"stars":null,"state":"unrated"},"availabilityStatus":"unknown","priceObserved":null,"sessionDeposit":null,"pricing":{"kind":"static","summary":"$0.01/call","primary":{"kind":"static","protocol":"x402","network":"base","amountUsd":"0.01","per":"call","confidence":"exact"},"accepted":[{"kind":"static","protocol":"x402","network":"base","amountUsd":"0.01","per":"call","confidence":"exact"}]},"paymentMethods":[{"uid":"pm_ilxISCXSPs55t2ycKGEU4","protocol":"x402","methodType":"crypto","chain":"base","mode":"charge","costAmount":"0.01","costPer":"request","priority":0,"asset":"0x833589fCD6eDb6E08f4c7C32D4f71b54bdA02913","unit":"request","depositMicros":null,"planRef":null}],"brandName":null,"brandSlug":null,"brandBaseUrl":null,"brandDocsUrl":null,"whatItDoes":"Fetches any public URL (article or PDF) and returns clean Markdown with title, byline, site name, excerpt, and word count using Firefox Readability extraction.","exampleAgentPrompt":"Can you fetch this article at https://en.wikipedia.org/wiki/Artificial_intelligence and give me the clean markdown text along with the title and word count?","exampleUseCases":[{"title":"Research pipeline content ingestion","prompt":"Pull the full text of this research paper PDF at https://arxiv.org/pdf/2301.00234.pdf and convert it to markdown so I can feed it into my summarizer."},{"title":"News article extraction for LLM","prompt":"Grab the article at https://www.nytimes.com/2024/01/15/technology/ai-regulation.html and return the clean markdown, the byline, and how many words it is."},{"title":"Wikipedia page to structured text","prompt":"Convert the Wikipedia page at https://en.wikipedia.org/wiki/Climate_change to clean markdown — I need the title, site name, and the full article body without any HTML clutter."}],"resultDescription":"Returns a structured response containing: the full article or PDF content as clean Markdown text, the page title, byline/author, site name, a short excerpt, and word count (or page count for PDFs). For image-only PDFs, returns a no_text_layer status. For JavaScript-rendered app shells that cannot be extracted, returns not_extractable — never silently returns an empty or misleading result.","failureModes":["Image-only PDFs return no_text_layer status with no markdown body","Client-rendered single-page apps return not_extractable when no static HTML is available","Private or internal hostnames are refused with an error","URLs exceeding the 8 MB fetch cap or 400K character output cap return an error","Paywalled or login-protected pages may return incomplete or no content","Invalid or malformed URLs return an error"],"whenToPreferThis":"Choose this endpoint when you need reliable, honest article or PDF text extraction from a public URL with clean Markdown output and rich metadata (title, byline, excerpt, word count). It is ideal for AI research pipelines, summarizers, and content ingestion workflows that need to trust the output — unlike generic scrapers that silently return empty or partial pages, this endpoint explicitly signals failures like image-only PDFs or non-extractable app shells. Prefer it over raw HTML fetchers when your downstream task is text-based (LLM summarization, RAG ingestion, content analysis).","instructions":null,"reviewSummary":null,"reviewSummaryHighlights":null,"reviewSummaryConcerns":null,"reviewSummaryGeneratedAt":null,"activationCount":0,"lastUsedAt":null,"lastSuccessfullyRanAt":null,"lastHealthCheckAt":"2026-10-01T12:51:38.983Z","isFirstParty":false,"canonicalSlug":"web-article-pdf-to-markdown-converter-fe561683"}