{"uid":"cap_VuvUH8GBBtvzvCpw1Zyvf","slug":"llama-api-pay-per-call-chat-completions-dd28e0f4","name":"Llama API (Pay-Per-Call Chat Completions)","description":"Llama API - pay per call, no API key. OpenAI-format chat completions on meta-llama/llama-4-maverick, pinned to DeepInfra for reliability. You sign a fixed price computed from your max_tokens; unused tokens are not refunded. 20s server-side timeout - never charged if it fires; set your client timeout to 30s. Try GET /llm/llama/sample.","url":"https://x402.agentindex.world/llm/llama?utm_source=zero.xyz","method":"GET","headers":{},"bodySchema":{"type":"object","properties":{"messages":{"type":"array","items":{"type":"object"},"minItems":1,"description":"OpenAI-format chat messages: [{role, content}, ...]."},"max_tokens":{"type":"integer","description":"Max completion tokens, capped at 4096 - also bounds the price ceiling."}}},"responseSchema":null,"example":null,"exampleRequest":null,"tags":["x402"],"displayCostAmount":"0.001","displayCostAsset":"USDC","priceDynamic":false,"priceHint":null,"priceStatus":"priced","priceSource":"probe","requiresHandshake":false,"reviewCount":0,"rating":{"score":"0.00","successRate":"0.00","reviews":0,"stars":null,"state":"unrated"},"availabilityStatus":"unknown","priceObserved":null,"sessionDeposit":null,"pricing":{"kind":"static","summary":"$0.001/call","primary":{"kind":"static","protocol":"x402","network":"base","amountUsd":"0.001","per":"call","confidence":"exact"},"accepted":[{"kind":"static","protocol":"x402","network":"base","amountUsd":"0.001","per":"call","confidence":"exact"}]},"paymentMethods":[{"uid":"pm_9njuR7PmR5LjgbBi9N4Mg","protocol":"x402","methodType":"crypto","chain":"base","mode":"charge","costAmount":"0.001","costPer":"request","priority":0,"asset":"0x833589fCD6eDb6E08f4c7C32D4f71b54bdA02913","unit":"request","depositMicros":null,"planRef":null}],"brandName":null,"brandSlug":null,"brandBaseUrl":null,"brandDocsUrl":null,"whatItDoes":"Pay-per-call OpenAI-format chat completions powered by Meta Llama 4 Maverick via DeepInfra, with no API key required and fixed per-call pricing.","exampleAgentPrompt":"Ask Llama 4 Maverick: 'Explain the difference between supervised and unsupervised learning in 3 short paragraphs' — use up to 512 tokens for the reply.","exampleUseCases":[{"title":"On-demand LLM for agent pipeline","prompt":"I need you to send this conversation history to Llama 4 Maverick and get a response: [{\"role\":\"user\",\"content\":\"Summarize the key risks of investing in meme coins in 5 bullet points.\"}] — limit the reply to 300 tokens."},{"title":"No-key prototype LLM call","prompt":"I want to test a quick LLM call without setting up an API key — send this message to Llama 4 and show me what it says: 'Write a haiku about distributed systems.' Keep max_tokens at 100."},{"title":"Chatbot response generation","prompt":"My chatbot needs a reply for a user who asked 'What are three healthy breakfast ideas for someone with no time in the morning?' — run it through Llama 4 Maverick with a 400 token limit and return the response."}],"resultDescription":"An OpenAI-compatible chat completion response containing the assistant's generated text, produced by Meta Llama 4 Maverick. The response reflects the conversation history provided in the messages array and is bounded by the max_tokens parameter. Unused tokens within the ceiling are not refunded. The server enforces a 20-second timeout and charges are not applied if the timeout fires.","failureModes":["20-second server-side timeout fires — request is not charged but no completion is returned","max_tokens exceeds 4096 cap — request may be rejected or capped","messages array is empty or malformed — 400-level error returned","USDC payment fails or is insufficient — payment-gated 402 response, no completion returned","Client timeout shorter than 30s causes premature disconnect before server responds"],"whenToPreferThis":"Choose this endpoint when you need a pay-per-call LLM completion with no API key setup, no subscription, and predictable per-call USDC pricing. It is ideal for AI agents that need sporadic or burst LLM calls without committing to a provider account, or for prototyping pipelines using the OpenAI message format against a capable open-weights model (Llama 4 Maverick). Prefer alternatives if you need streaming responses, model selection flexibility, refunds for unused tokens, or latency under a few seconds consistently.","instructions":null,"reviewSummary":null,"reviewSummaryHighlights":null,"reviewSummaryConcerns":null,"reviewSummaryGeneratedAt":null,"activationCount":0,"lastUsedAt":null,"lastSuccessfullyRanAt":null,"lastHealthCheckAt":"2026-10-01T06:40:36.050Z","isFirstParty":false,"canonicalSlug":"llama-api-pay-per-call-chat-completions-dd28e0f4"}