{"uid":"cap_TlvuIHFyJBaKGPZ6tmQi7","slug":"agent402-tools-pro-chat-completions-175d6487","name":"agent402.tools Pro Chat Completions","description":"OpenAI-compatible chat completions, pro tier: gpt-4o, gpt-4.1, claude sonnet, gemini pro, grok - paid per call in USDC over x402. Same wire format as /v1/chat/completions with higher input/output caps (48k chars in, 4096 tokens out).","url":"https://agent402.tools/v1/pro/chat/completions","method":"POST","headers":{},"bodySchema":{"type":"object","properties":{"zdr":{"type":"boolean","description":"Optional - true routes only to zero-data-retention providers (OpenRouter provider.zdr); the only provider preference a caller may set."},"model":{"type":"string","description":"Model id - OpenRouter form (openai/gpt-4o-mini) or bare OpenAI form (gpt-4o-mini). GET /v1/models lists the allowlist per tier. Optional: omit it and the tier serves its documented default (x402.defaultModel on /v1/models), named back in agent402_default_model; the price does not change"},"tools":{"type":"array","description":"Optional - OpenAI function tools {type:\"function\", function:{...}}, or a tool namespace {type:\"namespace\", name, tools:[...]} (flattened into its functions). The pro and premium routes also accept the bounded server tools openrouter:web_search, openrouter:web_fetch and openrouter:datetime with server-owned limits (GET /v1/models lists them); stop_server_tools_when and max_tool_calls are refused. A request with a server tool is never served from the prompt cache."},"messages":{"type":"array","description":"OpenAI chat messages: [{role, content}] - text and image_url content blocks supported"},"reasoning":{"type":"object","description":"Optional - {effort: \"none\"|\"minimal\"|\"low\"|\"medium\"|\"high\"|\"xhigh\"|\"max\", max_tokens?, exclude?, enabled?}. Reasoning tokens count against max_tokens. Omitted: low effort on the budget tiers, the model default on premium. reasoning_effort (string) is accepted as an alias."},"max_tokens":{"type":"number","description":"Output token cap (clamped to the tier maximum)"},"cache_control":{"description":"Optional - prompt caching preference. Default ON ({type:\"ephemeral\"}, 5-minute TTL): repeated prefixes across your turns are served from the provider cache (same price to you). Send false to disable. ttl:\"1h\" is not offered."},"max_completion_tokens":{"type":"integer","description":"Optional - alias of max_tokens (newer OpenAI SDKs send this)."}}},"responseSchema":{"type":"json","example":{"id":"gen-…","model":"openai/gpt-4o","usage":{"total_tokens":13,"prompt_tokens":12,"completion_tokens":1},"object":"chat.completion","choices":[{"index":0,"message":{"role":"assistant","content":"OK"},"finish_reason":"stop"}],"created":1750000000}},"example":null,"exampleRequest":null,"tags":["x402"],"displayCostAmount":"0.1","displayCostAsset":"USDC","priceDynamic":false,"priceHint":null,"priceStatus":"priced","priceSource":"probe","requiresHandshake":false,"reviewCount":0,"rating":{"score":"0.00","successRate":"0.00","reviews":0,"stars":null,"state":"unrated"},"availabilityStatus":"unknown","priceObserved":null,"sessionDeposit":null,"pricing":{"kind":"static","summary":"$0.1/call","primary":{"kind":"static","protocol":"x402","network":"base","amountUsd":"0.1","per":"call","confidence":"exact"},"accepted":[{"kind":"static","protocol":"x402","network":"base","amountUsd":"0.1","per":"call","confidence":"exact"}]},"paymentMethods":[{"uid":"pm_XQW0P88a1K7bn3ECuuFd8","protocol":"x402","methodType":"crypto","chain":"base","mode":"charge","costAmount":"0.1","costPer":"request","priority":0,"asset":"0x833589fCD6eDb6E08f4c7C32D4f71b54bdA02913","unit":"request","depositMicros":null,"planRef":null}],"brandName":null,"brandSlug":null,"brandBaseUrl":null,"brandDocsUrl":null,"whatItDoes":"OpenAI-compatible chat completions at pro tier supporting GPT-4o, GPT-4.1, Claude Sonnet, Gemini Pro, and Grok with higher input/output caps, paid per call in USDC via x402","exampleAgentPrompt":"Send this 40,000-character research document to Claude Sonnet via the pro chat completions endpoint and ask it to summarize the key findings — pay per call in USDC, and make sure zero-data-retention mode is on.","exampleUseCases":null,"resultDescription":"An OpenAI-compatible chat completion response object containing the assistant's generated message, finish reason, and token usage stats. Output is capped at 4096 tokens with up to 48k characters of input accepted. The response follows the same wire format as /v1/chat/completions so it works with any OpenAI SDK.","failureModes":["Payment failure: x402 payment not accepted or wallet has insufficient USDC balance","Model not available: requested model not on the pro-tier allowlist (check GET /v1/models)","ZDR routing failure: zdr=true requested but no zero-data-retention provider available for that model, walks failover chain and may error","Input too large: input exceeds 48k character cap","Max tokens exceeded: max_tokens clamped to tier maximum (4096)","Rate limiting or upstream provider outage causing 5xx errors","Invalid message format: messages array malformed or unsupported content block type"],"whenToPreferThis":"Use this endpoint when you need access to top-tier models (GPT-4o, GPT-4.1, Claude Sonnet, Gemini Pro, Grok) via a single OpenAI-compatible interface, especially when paying per call in USDC is preferred over a subscription. Ideal for agents that need large context windows (up to 48k chars input), want to avoid managing multiple API keys, require zero-data-retention routing for sensitive workloads, or want crypto-native pay-as-you-go LLM inference. Choose over the standard tier when you need the higher-capability models or larger I/O caps.","instructions":null,"reviewSummary":null,"reviewSummaryHighlights":null,"reviewSummaryConcerns":null,"reviewSummaryGeneratedAt":null,"activationCount":0,"lastUsedAt":null,"lastSuccessfullyRanAt":null,"lastHealthCheckAt":"2026-09-13T18:51:55.722Z","isFirstParty":false}