{"uid":"cap_oF4vm91WJCaqrdyNkyAb4","slug":"x402-llm-gateway-extended-tier-32k-output-78c881de","name":"x402 LLM Gateway – Extended Tier (32K output)","description":"Extended tier of the x402 LLM gateway: same self-hosted 27B open-weight LLM with the highest output ceiling (32,768 tokens) for long-form writing, reports and full document drafts. Pay per call in USDC on Base via x402 - no API key or account. OpenAI-compatible chat completions. Pick the cheaper quick/long tiers for shorter work.","url":"https://api.erb-llm.com/v1/extended","method":"POST","headers":{},"bodySchema":{"type":"object","properties":{"model":{"type":"string","description":"Model ID from the free GET /v1/models endpoint (e.g. qwen/qwen3.8-27b). Optional - a default chat model is chosen if omitted."},"messages":{"type":"array","items":{"type":"object","required":["role","content"],"properties":{"role":{"enum":["system","user","assistant"],"type":"string"},"content":{"type":"string"}}},"minItems":1,"description":"OpenAI chat messages (system/user/assistant)."},"max_tokens":{"type":"integer","minimum":1,"description":"Optional output token ceiling. This tier caps output at 32768 tokens - higher values are clamped to 32768, not rejected. Use the long/extended tiers for bigger outputs."},"temperature":{"type":"number","maximum":2,"minimum":0}}},"responseSchema":null,"example":null,"exampleRequest":null,"tags":["x402"],"displayCostAmount":"0.2","displayCostAsset":"USDC","priceDynamic":false,"priceHint":null,"priceStatus":"priced","priceSource":"probe","requiresHandshake":false,"reviewCount":0,"rating":{"score":"0.00","successRate":"0.00","reviews":0,"stars":null,"state":"unrated"},"availabilityStatus":"unknown","priceObserved":null,"sessionDeposit":null,"pricing":{"kind":"static","summary":"$0.2/call","primary":{"kind":"static","protocol":"x402","network":"base","amountUsd":"0.2","per":"call","confidence":"exact"},"accepted":[{"kind":"static","protocol":"x402","network":"base","amountUsd":"0.2","per":"call","confidence":"exact"}]},"paymentMethods":[{"uid":"pm_0pUd7Z7nybnZTHjpR_OXN","protocol":"x402","methodType":"crypto","chain":"base","mode":"charge","costAmount":"0.2","costPer":"request","priority":0,"asset":"0x833589fCD6eDb6E08f4c7C32D4f71b54bdA02913","unit":"request","depositMicros":null,"planRef":null}],"brandName":null,"brandSlug":null,"brandBaseUrl":null,"brandDocsUrl":null,"whatItDoes":"Serves OpenAI-compatible chat completions from a self-hosted 27B open-weight LLM with a 32,768-token output ceiling, billed at $0.20 USDC per call via the x402 protocol on Base — no API key required.","exampleAgentPrompt":"Write me a full 20-page technical report on the state of quantum computing in 2024 — use up to 32,000 tokens, set temperature to 0.7, and use the default model.","exampleUseCases":[{"title":"Full technical report generation","prompt":"Draft a comprehensive 15,000-word market analysis report on the electric vehicle industry in Europe for 2024, covering regulatory trends, key players, and consumer adoption curves. Use the extended LLM tier and set temperature to 0.4 so it stays factual."},{"title":"Long-form fiction writing","prompt":"Write me a full short-story collection — three interconnected stories totaling around 25,000 words — set in a near-future dystopian city. Use the 27B open-weight model on the extended tier and keep temperature at 0.9 for creativity."},{"title":"Detailed legal document drafting","prompt":"Draft a thorough software licensing agreement covering SaaS usage, IP ownership, liability caps, and GDPR compliance clauses — I need the full document, so use the extended tier with up to 32,768 tokens and a temperature of 0.2 to keep it precise."}],"resultDescription":"A standard OpenAI-format chat completion response object containing the assistant's generated message. The output can be up to 32,768 tokens long. Includes the model used, finish reason, and token usage counts. Suitable for long-form documents, detailed reports, and multi-section content.","failureModes":["Payment failure: x402 payment not accepted or insufficient USDC balance on Base — request rejected before processing","Invalid messages array: missing required 'role' or 'content' fields returns a 400 validation error","Model not found: specifying an invalid model ID returns a 404 or model-not-found error","Token clamping: max_tokens values above 32,768 are silently clamped to 32,768, not rejected — output may be shorter than expected","Rate limiting or capacity issues on the self-hosted infrastructure may return 503 or timeout errors","Very long prompts consuming most of the context window may leave little room for output generation"],"whenToPreferThis":"Choose this endpoint when you need the maximum possible output length (up to 32,768 tokens) in a single call — such as full document drafts, long reports, or complete multi-section content. Prefer it over the quick or long tiers only when output length justifies the $0.20/call price. Ideal for agents that need no API key management and can pay per call in USDC on Base. Use the cheaper quick tier for short responses and the long tier for mid-length outputs.","instructions":null,"reviewSummary":null,"reviewSummaryHighlights":null,"reviewSummaryConcerns":null,"reviewSummaryGeneratedAt":null,"activationCount":0,"lastUsedAt":null,"lastSuccessfullyRanAt":null,"lastHealthCheckAt":"2026-09-15T12:31:32.549Z","isFirstParty":false}