{"uid":"cap_-5VelUQ9Ea4GkRa8SlNal","slug":"dicta-notes-long-form-transcription-061ef250","name":"Dicta Notes Long-Form Transcription","description":"Turn a full recording into a structured document: POST JSON with an audio_url — up to 2 HOURS of meeting, interview, hearing, or podcast audio — and get a diarized who-said-what transcript JSON: timestamps, speaker tracking with per-segment confidence, language codes. Most transcription APIs cap at ~20 minutes; this one runs Gemini long-context, built for the long ones. mp3, wav, m4a, ogg, webm, flac. Flat price, USDC on Base, no account.","url":"https://dicta-notes.com/routes/x402/transcribe-long","method":"POST","headers":{},"bodySchema":{"type":"object","properties":{"store":{"type":"boolean","description":"Optional. When true, encrypt and store the transcript JSON for 60 days."},"audio_url":{"type":"string","format":"uri","description":"Public http(s) URL of the audio file (up to 2 hours, 400 MB max)."}}},"responseSchema":null,"example":null,"exampleRequest":null,"tags":["x402"],"displayCostAmount":"0.59","displayCostAsset":"USDC","priceDynamic":false,"priceHint":null,"priceStatus":"priced","priceSource":"probe","requiresHandshake":false,"reviewCount":0,"rating":{"score":"0.00","successRate":"0.00","reviews":0,"stars":null,"state":"unrated"},"availabilityStatus":"unknown","priceObserved":null,"sessionDeposit":null,"pricing":{"kind":"static","summary":"$0.59/call","primary":{"kind":"static","protocol":"x402","network":"base","amountUsd":"0.59","per":"call","confidence":"exact"},"accepted":[{"kind":"static","protocol":"x402","network":"base","amountUsd":"0.59","per":"call","confidence":"exact"}]},"paymentMethods":[{"uid":"pm_KLEZmotuTI_MxMvLI_-ND","protocol":"x402","methodType":"crypto","chain":"base","mode":"charge","costAmount":"0.59","costPer":"request","priority":0,"asset":"0x833589fCD6eDb6E08f4c7C32D4f71b54bdA02913","unit":"request","depositMicros":null,"planRef":null}],"brandName":null,"brandSlug":null,"brandBaseUrl":null,"brandDocsUrl":null,"whatItDoes":"Transcribes up to 2 hours of audio into a diarized, timestamped JSON transcript with per-speaker segment tracking and confidence scores, powered by Gemini long-context","exampleAgentPrompt":"Can you transcribe this 90-minute board meeting recording for me — I need a full speaker-by-speaker breakdown with timestamps, showing exactly who said what throughout the whole thing? Here's the audio URL: https://example.com/board-meeting-jan2025.mp3","exampleUseCases":[{"title":"Post-meeting transcript with speaker labels","prompt":"I have a 2-hour Zoom recording of our all-hands meeting at audio URL https://storage.example.com/allhands-feb2025.m4a — can you transcribe it and give me a full breakdown of who said what with timestamps so I can share notes with the team?"},{"title":"Journalistic interview diarization","prompt":"I recorded a 75-minute interview with three sources for my article — the audio is at https://recordings.press/interview-march.wav. Can you transcribe it and label each speaker's segments with timestamps and confidence scores so I can quickly find each person's quotes?"},{"title":"Legal hearing transcript generation","prompt":"We have a 90-minute court hearing audio file at https://legalfiles.firm.com/hearing-2025-03-15.mp3 — can you produce a structured JSON transcript with speaker diarization and timestamps so we can review who said what during the proceedings?"}],"resultDescription":"A JSON object containing a diarized transcript broken into segments, each with a speaker identifier, start and end timestamps, spoken text, per-segment confidence score, and detected language code — covering the full audio duration up to 2 hours.","failureModes":["Audio URL unreachable or returns non-audio content — transcription fails with an error","Audio format not supported (must be mp3, wav, m4a, ogg, webm, or flac)","Audio exceeds 2-hour limit — request rejected or truncated","Poor audio quality or heavy background noise reduces transcription accuracy and confidence scores","Insufficient USDC balance or payment failure on Base — request not processed","Speaker diarization may merge or split speakers in noisy or overlapping-speech segments"],"whenToPreferThis":"Choose this endpoint when you need to transcribe audio longer than ~20 minutes — the typical cap for most transcription APIs — with speaker diarization included. It is specifically built for long-form content: multi-hour meetings, interviews, hearings, and podcasts where knowing who said what at each timestamp is as important as the words themselves. Prefer this over short-audio endpoints when the recording exceeds 20 minutes or when per-speaker segmentation with confidence scores is required. Payment is flat-rate USDC on Base with no account needed.","instructions":null,"reviewSummary":null,"reviewSummaryHighlights":null,"reviewSummaryConcerns":null,"reviewSummaryGeneratedAt":null,"activationCount":0,"lastUsedAt":null,"lastSuccessfullyRanAt":null,"lastHealthCheckAt":"2026-09-13T12:32:47.349Z","isFirstParty":false}