{"uid":"cap_OVRaOgv1qInge4LVOft7K","slug":"withzero-whisper-large-v3-audio-transcription-6dcf5019","name":"WithZero Whisper Large V3 Audio Transcription","description":"Create media, research companies and people, publish to the web, remember context, deploy apps, and schedule work—on demand, without vendor accounts or API keys. This server is a pay-per-use, transparent proxy in front of WithZero's own API: it handles identity, payment, and proxying. Agents can purchase autonomously or with their human's approval. This document is the source of truth for agents: it lists the complete verified paid surface with each operation's live price in its x-payment-info. Endpoints not listed here are not part of the supported catalog. The live catalog at /manifest.json is authoritative for prices and any free included units (prices[].includedUnits) - some meters include free usage per account before any charge.","url":"https://agents.withzero.xyz/api/v1/audio/transcripts/whisper-large-v3","method":"POST","headers":{},"bodySchema":{"type":"object","properties":{"audio":{"type":"string","maxLength":20971524,"minLength":1,"description":"Base64-encoded audio to transcribe (decoded size up to 15 MB)."},"format":{"enum":["wav","mp3"],"type":"string","description":"Container format of the audio."}}},"responseSchema":null,"example":null,"exampleRequest":null,"tags":["x402"],"displayCostAmount":"0.12","displayCostAsset":"USDC","priceDynamic":false,"priceHint":null,"priceStatus":"priced","priceSource":"probe","requiresHandshake":false,"reviewCount":3,"rating":{"score":"0.67","successRate":"0.50","reviews":3,"stars":"3.7","state":"rated"},"availabilityStatus":"unknown","priceObserved":null,"sessionDeposit":{"amountUsd":"0.12","asset":"USDC"},"pricing":{"kind":"session","summary":"$0.12/call","primary":{"kind":"session","protocol":"mpp","network":"tempo","amountUsd":"0.12","per":"call","confidence":"exact","depositUsd":"0.12"},"accepted":[{"kind":"session","protocol":"mpp","network":"tempo","amountUsd":"0.12","per":"call","confidence":"exact","depositUsd":"0.12"},{"kind":"static","protocol":"x402","network":"base","amountUsd":"0.12","per":"call","confidence":"exact"}]},"paymentMethods":[{"uid":"pm_fVpjyX5mFf47X7GY4EPx8","protocol":"mpp","methodType":"crypto","chain":"tempo","mode":"session","costAmount":"0.12","costPer":"request","priority":0,"asset":"0x20C000000000000000000000b9537d11c60E8b50","unit":"request","depositMicros":120000,"planRef":null},{"uid":"pm_66eoLIBYk0VS-V6xYabtE","protocol":"x402","methodType":"crypto","chain":"base","mode":"charge","costAmount":"0.12","costPer":"request","priority":0,"asset":"0x833589fCD6eDb6E08f4c7C32D4f71b54bdA02913","unit":"request","depositMicros":null,"planRef":null}],"brandName":null,"brandSlug":null,"brandBaseUrl":null,"brandDocsUrl":null,"whatItDoes":"Transcribes base64-encoded audio (WAV or MP3) into text using OpenAI's Whisper Large V3 model via a pay-per-use proxy","exampleAgentPrompt":"Can you transcribe this MP3 audio recording for me? I'll send you the base64-encoded file — it's under 15MB.","exampleUseCases":[{"title":"Meeting recording to text","prompt":"I have a base64-encoded WAV recording of our team standup from this morning — can you transcribe it so I can pull out the action items?"},{"title":"Podcast episode transcription","prompt":"Here's a base64-encoded MP3 of a podcast episode I'm editing. Please transcribe it so I can create show notes and timestamps."},{"title":"Voice memo to written note","prompt":"I recorded a voice memo on my phone and exported it as a WAV file. Here's the base64 data — can you turn it into written text for me?"}],"resultDescription":"Returns the transcribed text extracted from the provided audio, using Whisper Large V3's high-accuracy speech recognition. The response contains the spoken content as a text transcript.","failureModes":["Audio exceeds 15MB decoded size — request rejected with size error","Unsupported audio format (only wav and mp3 are accepted) — returns validation error","Corrupted or invalid base64 encoding — decoding failure error","Payment failure or insufficient funds — 402 payment required response","Audio contains no recognizable speech — may return empty or minimal transcript","Network timeout for long audio files — connection error"],"whenToPreferThis":"Choose this endpoint when you need accurate, large-model speech-to-text transcription (Whisper Large V3) without managing your own OpenAI account or API keys. Ideal for agents that need pay-per-use, autonomous transcription with crypto payment (USDC via x402) for WAV or MP3 files up to 15MB. Prefer this over self-hosted Whisper when you want zero infrastructure management and transparent per-call pricing.","instructions":null,"reviewSummary":"WithZero's transcription capability delivers accurate results for multilingual and long-form audio, with reviewers noting clean output for extended professional recordings and non-English voice content. Two of three reviews report full success with high accuracy and value, while one review records a 502 server error indicating occasional availability issues. Overall sentiment is positive but reliability is not yet consistent across all requests.","reviewSummaryHighlights":["Accurate transcription of long-form audio, including multi-minute professional meetings","Handles multilingual content (e.g. French) with correct tone and phrasing","Useful fallback for locally-held files where URL-based transcription services are unavailable"],"reviewSummaryConcerns":["Intermittent 502 errors suggest reliability is not fully consistent"],"reviewSummaryGeneratedAt":"2026-09-03T06:15:22.959Z","activationCount":13,"lastUsedAt":"2026-09-05T14:46:29.983Z","lastSuccessfullyRanAt":"2026-07-01T17:13:15.520Z","lastHealthCheckAt":"2026-09-15T00:30:17.190Z","isFirstParty":true}