{"uid":"cap_x6SC3Vxqp1a0P-f1WySgq","slug":"modelprices-xyz-cheapest-vision-multimodal-llm-leaderboard-5643e779","name":"modelprices.xyz Cheapest Vision/Multimodal LLM Leaderboard","description":"Cheapest vision/multimodal LLM leaderboard: the 50 lowest-cost AI models that accept image input, ranked by token price across every provider — GPT-5, Claude 5, Gemini 3, Llama 4, Qwen VL and more. Input, output and cache USD per 1M tokens with context window joined in. Answers 'what is the cheapest model that can read images?' in one call. Refreshed hourly.","url":"https://modelprices.xyz/llm/cheapest/vision","method":"GET","headers":{},"bodySchema":{"type":"object","$schema":"https://json-schema.org/draft/2020-12/schema","required":["input"],"properties":{"input":{"type":"object","required":["type","method"],"properties":{"type":{"type":"string","const":"http"},"method":{"enum":["GET"],"type":"string"},"queryParams":{"type":"object","properties":{}}},"additionalProperties":false},"output":{"type":"object","required":["type"],"properties":{"type":{"type":"string"},"example":{"type":"object"}}}}},"responseSchema":null,"example":null,"exampleRequest":null,"tags":["x402"],"displayCostAmount":"0.01","displayCostAsset":"USDC","priceDynamic":false,"priceHint":null,"priceStatus":"priced","priceSource":"probe","requiresHandshake":false,"reviewCount":0,"rating":{"score":"0.00","successRate":"0.00","reviews":0,"stars":null,"state":"unrated"},"availabilityStatus":"unknown","priceObserved":null,"sessionDeposit":null,"pricing":{"kind":"static","summary":"$0.01/call","primary":{"kind":"static","protocol":"x402","network":"base","amountUsd":"0.01","per":"call","confidence":"exact"},"accepted":[{"kind":"static","protocol":"x402","network":"base","amountUsd":"0.01","per":"call","confidence":"exact"}]},"paymentMethods":[{"uid":"pm_yqYe8bICLvY-MyB7C_MLW","protocol":"x402","methodType":"crypto","chain":"base","mode":"charge","costAmount":"0.01","costPer":"request","priority":0,"asset":"0x833589fCD6eDb6E08f4c7C32D4f71b54bdA02913","unit":"request","depositMicros":null,"planRef":null}],"brandName":null,"brandSlug":null,"brandBaseUrl":null,"brandDocsUrl":null,"whatItDoes":"Returns the 50 lowest-cost AI models that accept image input, ranked by token price across all providers, with input, output, and cache pricing per 1M tokens.","exampleAgentPrompt":"What's the cheapest AI model right now that can accept image inputs? Pull the full vision model price leaderboard so I can compare input and output token costs across providers like OpenAI, Anthropic, Google, and Meta.","exampleUseCases":[{"title":"Cheapest image-capable model for a startup","prompt":"I'm building a product that needs to analyze product photos and I want to keep costs low — can you pull the cheapest vision LLM leaderboard and tell me which model gives me image input at the lowest price per million tokens?"},{"title":"Comparing multimodal model costs across providers","prompt":"I need to decide between GPT-5, Claude 5, and Gemini 3 for a vision task — can you fetch the current cheapest vision model rankings so I can see how their input and output token prices compare?"},{"title":"Selecting a budget vision model for high-volume inference","prompt":"I'm going to be running thousands of image-analysis calls per day and budget is tight — what are the top 10 lowest-cost AI models that support image input right now, including cache pricing?"}],"resultDescription":"A ranked list of up to 50 vision/multimodal LLM models, each with model name, provider, input cost per 1M tokens, output cost per 1M tokens, cache cost per 1M tokens, and context window size — sorted ascending by price and refreshed hourly.","failureModes":["Upstream pricing data unavailable — service may return stale or partial data","HTTP 402 if payment is not completed via x402 protocol","No models returned if the filter criteria exclude all vision-capable models","Hourly refresh delay means prices could be up to 60 minutes behind actual provider changes"],"whenToPreferThis":"Use this endpoint when you specifically need to find or compare AI models that support image/vision input ranked by cost. Prefer this over the general cheapest LLM leaderboard when the image-processing capability is a hard requirement. Use this instead of individual provider pricing tables (e.g. OpenAI or Bedrock) when you want a cross-provider comparison in a single call.","instructions":null,"reviewSummary":null,"reviewSummaryHighlights":null,"reviewSummaryConcerns":null,"reviewSummaryGeneratedAt":null,"activationCount":0,"lastUsedAt":null,"lastSuccessfullyRanAt":null,"lastHealthCheckAt":"2026-09-15T00:50:22.994Z","isFirstParty":false}