AI model overview
See which platforms expose which models, what they cost through the main public route, and where the wrapper/platform choice matters more than the model name.
Updated August 2026: model names and prices moved fast
Sora is now legacy context, OpenAI image work should be read as GPT Image rather than DALL-E, Veo is the clearest native-audio video lane, and Ideogram plus Flux now deserve more weight in image workflows. Pricing is shown by access route because the same model can be sold as API usage, plan credits, bundled subscription access, or self-hosted infrastructure.
Start with the live buyer guides feeding this page
If you arrived here from a model question, anchor it in a real buying workflow first. For the live analyst branch, keep the order literal: use /compare first when the packaged shortlist is down to real finalists, keep the PDF-analysis guide as the active feeder that pressure-tests those document-heavy finalists, use the broader data-analysis guide only as the wider support branch, and open /stack only after that when the blocker is which model access route or wrapper depth actually differs. This branch should stay separate from the colder note-taking and research detour on /recommend.
Fully open-source image generation model offering maximum customization and fine-tuning capabilities.
Stable Diffusion 3.5 pricing depends on whether you self-host or use a managed endpoint.
Source · verified 2026-08-06Black Forest Labs' image generation and editing model family for text-to-image, image-to-image, and instruction-based edits with strong character and style consistency.
BFL also sells service credits directly; wrapper prices vary by provider and resolution.
Source · verified 2026-08-06Llama 4 Maverick is the previous-generation Meta workhorse and still one of the cheapest capable open-weight models on hosted APIs. It remains a common self-hosting baseline.
Weights are free; the quoted rate is a representative hosted price and varies by inference provider.
Source · verified 2026-08-13Llama 5 is Meta's 600B-parameter mixture-of-experts flagship with native video and audio input. Its 5M-token context window is the largest of any publicly available model, open or closed, and the weights are downloadable.
Released under the Llama 5 Community License, which is not OSI-approved. Review the commercial-use terms before deploying.
Source · verified 2026-08-13DeepSeek V4 pairs a 1M-token context window with open weights and unusually aggressive API pricing. It is the standard cost-sensitive pick when you want frontier-adjacent quality without frontier pricing.
1M context under an MIT license. Cache-hit input bills substantially lower than the cache-miss rate quoted here.
Source · verified 2026-08-13Kimi K2.6 is the multimodal agent tier, combining text, image, and video input with tool use at roughly a third of K3's price. Context caps at 256K, well below the K3 flagship.
Roughly a third of K3's cost, with multimodal input but a 256K context ceiling.
Source · verified 2026-08-13Kimi K3 is Moonshot's 2.8-trillion-parameter flagship and the largest open-weight model in the comparison set. It is priced like a frontier closed model rather than like the cheaper open-weight lane.
Priced like a closed frontier model despite shipping open weights — the largest open-weight model here at 2.8T parameters.
Source · verified 2026-08-13Mistral Small 4 is the Apache 2.0 workhorse of the Mistral line — small enough to self-host on modest hardware, cheap enough to run at volume, and unrestricted for commercial use.
Apache 2.0 weights with no commercial restrictions, and small enough to self-host on modest hardware.
Source · verified 2026-08-13Qwen3.5 Plus is the mid-tier Qwen model, trading some of the Flash line's cost advantage for stronger reasoning while keeping the 1M-token context window and Apache 2.0 license.
Mid-tier Qwen model, still Apache 2.0 and still a 1M-token context window.
Source · verified 2026-08-13Qwen3.7 Flash is the cheapest 1M-context API in the comparison set by a wide margin. Apache 2.0 licensing means the weights can also be self-hosted without a bespoke commercial agreement.
Cheapest per-token rate in this comparison set. Apache 2.0 weights also allow self-hosting.
Source · verified 2026-08-13DeepSeek V4 Flash is the low-latency V4 variant and one of the cheapest 1M-context models available anywhere, at a fraction of the flagship rates from OpenAI, Anthropic, and Google.
Quoted rate is for cache-miss input. One of the cheapest 1M-context APIs available.
Source · verified 2026-08-13GLM-5.2 is the long-context coding specialist of the open-weight field, consistently ranked against DeepSeek V4 and Qwen on repository-scale code tasks.
Positioned as the long-context coding pick among open-weight models. MIT-licensed weights.
Source · verified 2026-08-13GPT-5.6 Sol is OpenAI's flagship reasoning tier, the top of the three-tier GPT-5.6 family. It carries the largest published context window of any major closed model at 1.05M tokens.
Flagship reasoning tier. Output pricing is the highest of the mainstream closed flagships at this input rate.
Source · verified 2026-08-13Mistral Large 3 is the flagship of the European open-weight route, with EU data residency and on-prem deployment options that matter more to regulated buyers than raw benchmark position.
EU data residency and on-prem deployment are available, which is usually the deciding factor over raw price.
Source · verified 2026-08-13Claude Opus 5 is Anthropic's model for complex agentic coding and long-horizon enterprise work. Thinking is on by default, effort is tunable from low through max, and it ships a 1M-token context window at standard pricing.
1M context at standard pricing with no long-context premium. Prompt caching reads bill at roughly 0.1x.
Source · verified 2026-08-13Gemini 3.1 Pro is Google's flagship reasoning model. Pricing is tiered by prompt length: rates roughly double once a request passes 200K tokens, which matters more here than on flat-rate competitors.
Above 200K context the rate steps to $4.00 input / $18.00 output. Context caching cuts repeat-prompt cost by up to ~90%.
Source · verified 2026-08-13OpenAI's current state-of-the-art image generation and editing model, replacing GPT Image 1 as the safer default for new OpenAI image workflows.
Batch pricing is lower; image cost depends on generated token count, quality, and size.
Source · verified 2026-08-06GPT-5.6 Terra is the balanced production tier of the GPT-5.6 family, priced between Sol and Luna and intended as the general-purpose default for most API workloads.
Balanced production tier, priced level with Gemini 3.1 Pro's sub-200K rate.
Source · verified 2026-08-13Adobe's commercially oriented image model for brand-safe generation and Creative Cloud workflows.
Commercial buyers should compare Firefly credits inside the specific Adobe plan they already own.
Source · verified 2026-08-06Adobe's commercially oriented video generation model for Firefly text-to-video and image-to-video workflows, especially useful when Creative Cloud integration and IP-safe production controls matter.
Commercial buyers should compare included credits inside the Adobe plan they already own.
Source · verified 2026-08-13AI composer specializing in classical, cinematic, and orchestral music with MIDI export.
AIVA sells composition/download rights by plan rather than per-token generation.
Source · verified 2026-08-06Claude Fable 5 is Anthropic's most capable widely released model, priced above the Opus tier and aimed at the most demanding reasoning and long-horizon agentic work. Thinking is always on and cannot be disabled.
Priced above the Opus tier. Requires 30-day data retention and is not available under zero-data-retention terms.
Source · verified 2026-08-13Claude Haiku 4.5 is Anthropic's fastest and cheapest model, aimed at classification, extraction, and high-volume simple tasks. It is the only current Claude model that caps below 1M context.
Cheapest Claude tier. Context caps at 200K rather than the 1M offered by the Opus and Sonnet tiers.
Source · verified 2026-08-13Claude Sonnet 5 is the balanced Anthropic tier, reaching near-Opus quality on coding and agentic work at roughly half the price. Adaptive thinking is on by default and the context window matches the Opus tier at 1M tokens.
Introductory rate of $2.00 / $10.00 applies through 2026-08-31, after which standard pricing takes effect.
Source · verified 2026-08-13State-of-the-art text-to-speech and voice cloning model with 29 language support and emotional nuance.
Website generation uses credits; Eleven v3 is 1 credit per character in the app.
Source · verified 2026-08-06FLUX 3 Video is Black Forest Labs' multimodal video, audio, image, and action-prediction model line, adding a video-generation lane to a brand previously tracked on CompareGen only for images.
Keep separate from Flux image pricing because video and audio generation costs will not normalize to per-image rates.
Source · verified 2026-08-13Gemini 2.5 Pro is the previous-generation Google flagship, still available and still competitive on coding. It remains the cheapest Pro-tier entry point in the Gemini family.
Above 200K context the rate steps to $2.50 input / $15.00 output.
Source · verified 2026-08-13Gemini 3.6 Flash is the speed-first Gemini tier, positioned as the highest-throughput option in the family before 3.7 arrived. Pricing is promotional through the end of 2026.
Promotional pricing through 2026-12-31. Some third-party trackers still list the earlier $1.50 / $7.50 rate.
Source · verified 2026-08-13Gemini 3.7 Flash is Google's most capable Flash model, tuned for agentic workflows. Its listed rate is promotional through December 31, 2026 and increases in January 2027.
Promotional pricing through 2026-12-31; rates increase 2027-01-01. Budget accordingly on annual contracts.
Source · verified 2026-08-13Google's Gemini Omni Flash is a top current video-generation model on independent text-to-video, image-to-video, and video-editing leaderboards, and should be compared separately from the Veo product route.
Gemini Omni Flash is tracked separately because benchmark naming and access route are distinct from the Veo product page.
Source · verified 2026-08-13Runway's fourth-generation image-to-video model offering cinematic motion quality, prompt adherence, and visual fidelity for 5- or 10-second generations.
A 5-second Gen-4 generation is 60 credits; web plans bundle credits.
Source · verified 2026-08-06Runway's frontier video model for stronger motion quality, visual fidelity, and prompt adherence. Treat it as the premium Runway lane rather than a separate workflow category.
Runway's Max plan also frames 9,500 credits as 791s of Gen-4.5 or Gen-4.
Source · verified 2026-08-06GPT-5.4 is the previous-generation OpenAI mid-tier, still widely deployed and still available on the API. Input above 272K tokens is billed at double the standard rate.
Input above 272K tokens bills at 2x. Cached input receives a 75-90% discount.
Source · verified 2026-08-13GPT-5.6 Luna is the cost-optimized tier of the GPT-5.6 family, competing directly with Gemini Flash and the cheaper open-weight APIs on high-volume work.
Cost-optimized tier competing with Gemini Flash and the hosted open-weight APIs.
Source · verified 2026-08-13Grok 4.3 is xAI's value tier and its widest-context option at 1M tokens. Output pricing is notably cheaper than the 4.5/4.6 flagship line, which makes it the better fit for output-heavy work.
Cheapest xAI general-purpose tier and the only 1M-context option in the family.
Source · verified 2026-08-13Grok 4.6 is xAI's current flagship, aimed at long-running agents, coding, and research. Its 500K context is smaller than the Grok 4.3 line, and prompts past 200K tokens bill the entire request at double rate.
Once a prompt reaches 200K tokens the entire request bills at $4.00 input / $12.00 output. Cached input is $0.50 / 1M.
Source · verified 2026-08-13xAI's Grok Imagine Video 1.5 is the current Grok video lane for text or image to video generation, with strong leaderboard performance in image-to-video and native-audio comparisons.
Compare as subscription access unless xAI publishes a clean standalone video API price.
Source · verified 2026-08-13MiniMax's Hailuo video model focused on lifelike motion and emotion, with text and image to video generation.
Hailuo publishes many plan and model guides; verify the in-app checkout for exact current credits.
Source · verified 2026-08-06Ideogram's current visual intelligence model for photorealistic images, legible text, precise style control, editing, API workflows, MCP, and open-weight deployment.
Ideogram docs explicitly route current plan and API prices to the live pricing page.
Source · verified 2026-08-06Google's latest text-to-image model delivering stunning high-quality visuals with excellent precision and realism. Best-in-class for professional integration.
Google's current public image table prices image output by tokens and resolution.
Source · verified 2026-08-06Kuaishou's current Kling 3.0 model series improves consistency, photorealism, 3-15 second short-form output, native multilingual audio, and newer reference-control workflows.
Credit cost changes by model, mode, duration, and quality; route budgets through the live credit table.
Source · verified 2026-08-06Leonardo's foundational image model for prompt adherence, coherent text rendering, and iterative brand or concept-art workflows.
Phoenix pricing is effective token cost inside Leonardo rather than a simple public per-image rate.
Source · verified 2026-08-06LTX-2.5 is the current Lightricks open video model for synchronized audio-video generation, multi-shot scenes, footage editing, and local or API-backed production workflows.
LTX is a model-stack lane: cost depends on local hardware, cloud GPU, or the wrapper provider.
Source · verified 2026-08-13Luma's multimodal creative model for generating images and videos from text or reference images in one workflow.
Use the same Luma plan budget as Ray unless a specific API route is selected.
Source · verified 2026-08-06Midjourney's V7 image model, released in April 2025, with stronger prompt precision, richer texture detail, Draft Mode, and Omni Reference for consistent characters and objects.
Midjourney prices access as GPU time and relaxed/fast modes, not per image.
Source · verified 2026-08-06Midjourney Video V1 turns still images into 5-second image-to-video clips with optional motion prompting, making Midjourney relevant to creator-video comparisons beyond static image generation.
Video is image-to-video inside Midjourney's subscription workflow rather than a separate public API.
Source · verified 2026-08-13MiniMax H3 is an open-weights omni-modal video model that reads text, images, video, and audio as unified context and generates up to 15-second 2K clips with native stereo audio.
Use this as a directional MiniMax H3/Hailuo route, not a guaranteed API tariff.
Source · verified 2026-08-06Professional voiceover model with 120+ AI voices across 20 languages for e-learning and corporate content.
Murf budgets are best compared as included voiceover minutes per seat.
Source · verified 2026-08-06Meta's Muse Video is the media-generation sibling to Muse Image, previewed with exceptional visual fidelity and native audio support. Track it separately from Muse Spark, which is the reasoning and agent model.
Track Muse Video as model-stack context and keep it out of default buyer picks until access is clearer.
Source · verified 2026-08-13Pika's idea-to-video model focused on creative expression and quick content generation for social media.
Public per-second API pricing is not cleanly exposed, so compare effective plan credits.
Source · verified 2026-08-06World's first reasoning video model that can think, plan, and create studio-grade content. Native 1080p HD, 4x faster generation, and 3x cheaper per-second pricing.
Luma bundles Ray access into plan credits; API pricing is exposed through Luma API billing.
Source · verified 2026-08-06Recraft V3 is a design-focused image generation model for brand visuals, illustrations, and editable creative assets.
Design workflows should compare included generations, export rights, and team seats.
Source · verified 2026-08-06Real-time voice cloning model with cross-language localization and emotion control.
Resemble pricing depends on voice cloning, localization, API, and minute bundle.
Source · verified 2026-08-06Reve's image model is strongest when prompt adherence and detailed composition accuracy matter more than ecosystem breadth.
Treat as app-plan pricing until Reve exposes a stable public per-image API tariff.
Source · verified 2026-08-06ByteDance Seed's current Seedance model is built for 30-second audio-video storytelling, precise reference control, powerful editing, and multimodal text, image, audio, and video inputs.
Seedance access and pricing are less standardized than Runway or Vertex, so CompareGen should avoid a fake normalized price.
Source · verified 2026-08-13OpenAI's deprecated video and audio generation model. It still matters historically for narrative coherence and synchronized audio, but it is no longer a safe default for new production video workflows.
Keep as legacy context on CompareGen even though OpenAI still lists API prices.
Source · verified 2026-08-06Latest Suno model for AI music generation with improved vocal quality and longer song support.
Suno sells song generation as subscription credits, not a direct per-token API price.
Source · verified 2026-08-06Professional-grade AI music model producing near-indistinguishable audio quality with superior instrument separation.
Udio credits are consumed per generation set and plan quota rather than a single model tariff.
Source · verified 2026-08-06Google's flagship video model line with native audio, stronger prompt adherence, richer audiovisual quality, and 8-second 720p/1080p outputs in Flow, Gemini, and Vertex/AI Studio routes.
Veo 3.1 pricing varies by Fast/Lite tier, 720p/1080p/4K, and whether audio is generated.
Source · verified 2026-08-06Alibaba's Wan video family is the important open video-generation lane for teams that want text-to-video and image-to-video workflows with local or API-hosted access instead of a closed creative app.
Wan cost depends on whether the buyer self-hosts or uses fal/Replicate/other hosted routes.
Source · verified 2026-08-13Go back down-funnel from the model layer
Once you understand the stack, jump back into the workflow-first pages that actually turn model knowledge into platform decisions.
Workflow quiz
Answer four quick questions and get pointed to the right category and platform shortlist.
Side-by-side compare
Line up up to three platforms and check quality, value, and workflow fit before you commit.
Short-form video picks
Best options for Shorts, TikTok, Reels, and other fast-turn social clips.
UGC ad workflow picks
Use this when you need creator-style ads, avatar-led testimonials, and fast social testing loops.
Paid social creative picks
Choose the right platform for testing lots of ad variants without overspending on polish too early.
Social media ad generator picks
Exact-match guidance for buyers comparing AI social media ad makers, creative generators, and fast paid-social tools.