{"uid":"cap_1V6XNLhusFDy0oLXLN7Mk","slug":"modelprices-xyz-million-token-context-llm-leaderboard-ade3233d","name":"modelprices.xyz Million-Token-Context LLM Leaderboard","description":"Cheapest million-token-context LLM leaderboard: every AI model with a 1,000,000-token context window or larger, ranked by inference cost per token — Gemini 3 Pro, Gemini 3 Flash, Llama 4 Scout, GPT-5 long-context tiers and more. Input, output and cache USD per 1M tokens with exact context window and max output joined in. Answers 'what is the cheapest model that fits my whole corpus?' Refreshed hourly.","url":"https://modelprices.xyz/llm/cheapest/million-token-context","method":"GET","headers":{},"bodySchema":{"type":"object","$schema":"https://json-schema.org/draft/2020-12/schema","required":["input"],"properties":{"input":{"type":"object","required":["type","method"],"properties":{"type":{"type":"string","const":"http"},"method":{"enum":["GET"],"type":"string"},"queryParams":{"type":"object","properties":{}}},"additionalProperties":false},"output":{"type":"object","required":["type"],"properties":{"type":{"type":"string"},"example":{"type":"object"}}}}},"responseSchema":null,"example":null,"exampleRequest":null,"tags":["x402"],"displayCostAmount":"0.01","displayCostAsset":"USDC","priceDynamic":false,"priceHint":null,"priceStatus":"priced","priceSource":"probe","requiresHandshake":false,"reviewCount":0,"rating":{"score":"0.00","successRate":"0.00","reviews":0,"stars":null,"state":"unrated"},"availabilityStatus":"unknown","priceObserved":null,"sessionDeposit":null,"pricing":{"kind":"static","summary":"$0.01/call","primary":{"kind":"static","protocol":"x402","network":"base","amountUsd":"0.01","per":"call","confidence":"exact"},"accepted":[{"kind":"static","protocol":"x402","network":"base","amountUsd":"0.01","per":"call","confidence":"exact"}]},"paymentMethods":[{"uid":"pm_EJ5S2Y5SXiCPfK2HMnLIl","protocol":"x402","methodType":"crypto","chain":"base","mode":"charge","costAmount":"0.01","costPer":"request","priority":0,"asset":"0x833589fCD6eDb6E08f4c7C32D4f71b54bdA02913","unit":"request","depositMicros":null,"planRef":null}],"brandName":null,"brandSlug":null,"brandBaseUrl":null,"brandDocsUrl":null,"whatItDoes":"Returns a ranked list of every AI model with a 1M+ token context window, sorted by inference cost per token, with input/output/cache pricing in USD per 1M tokens.","exampleAgentPrompt":"Show me the cheapest AI models that support at least a 1 million token context window, ranked by cost per token — I want to see input, output, and cache pricing so I can pick the most affordable one for processing a huge document corpus.","exampleUseCases":[{"title":"Cheapest model for full codebase ingestion","prompt":"I need to feed my entire monorepo into an LLM for analysis — which million-token context models are cheapest right now? Show me input and output costs per million tokens so I can pick the most budget-friendly option."},{"title":"Comparing long-context Gemini vs GPT tiers","prompt":"Can you pull up the current leaderboard of models with 1M+ token context windows and tell me how Gemini 3 Pro and Flash compare to GPT-5 long-context tiers on cost per token?"},{"title":"Selecting model for nightly book-length summarization pipeline","prompt":"I'm building an automated pipeline that processes book-length documents every night — can you find me the absolute cheapest model with at least a million-token context window, including any caching discounts?"}],"resultDescription":"A ranked leaderboard of AI models with context windows of 1 million tokens or larger, each entry including model name, provider, exact context window size, max output tokens, and USD cost per 1M tokens for input, output, and cached tokens. Data is refreshed hourly.","failureModes":["Service unavailable or timeout if the hourly refresh is in progress","Empty result set if no models meeting the 1M-token threshold are indexed (unlikely but possible during data gaps)","Stale pricing data if the upstream provider APIs were unreachable during the last refresh cycle","HTTP 402 payment required if the x402 payment header is missing or the USDC balance is insufficient"],"whenToPreferThis":"Choose this endpoint when you specifically need to find and compare models with million-token-or-larger context windows ranked by cost — ideal for cost optimization of long-document, codebase, or large-corpus inference tasks. Prefer it over general LLM pricing tables when context window size is the primary constraint and you want pre-filtered, pre-ranked results rather than scanning all models manually.","instructions":null,"reviewSummary":null,"reviewSummaryHighlights":null,"reviewSummaryConcerns":null,"reviewSummaryGeneratedAt":null,"activationCount":0,"lastUsedAt":null,"lastSuccessfullyRanAt":null,"lastHealthCheckAt":"2026-09-15T00:48:06.098Z","isFirstParty":false}