Generative AI Model Ranking Matrix

Living benchmark, refreshed nightly

Frontier language, video, and image generation models compared side by side. Benchmarks, capabilities, release dates, and price, so you can pick a model for the job and not the hype.

Updated

Frontier LLMs: coding and reasoning

27 benchmarked models, newest first. Then 13 newly released models tagged live, found nightly and waiting on published benchmarks. Click a column to sort. Click a model for its full profile.

License Caps / filter 1 to 3 tabs

Swipe sideways for benchmarks and price

Model Vendor Params (B) License Capabilities Released Context SWE-Bench Verified SWE-Bench Pro Terminal-Bench Reasoning* AA Index $/M in Best for
Capabilities V Vision (image input) A Audio T Tools / function calling R Extended reasoning C Computer / browser use
column leader Quieter numerals are lower tier *Reasoning is a composite score (GPQA Diamond, HLE, AIME, ARC-AGI-2, normalised 0 to 100, directional).
What the columns mean
SWE-Bench Verified
Share of 500 real GitHub issues the model fixed end to end, tests passing. The best single signal for coding agents.
SWE-Bench Pro
A harder, contamination-resistant set of long-horizon coding tasks. Scores run lower than Verified for every model.
Terminal-Bench
Tasks completed in a real shell: builds, data wrangling, debugging. Measures agentic work, not chat.
Reasoning
Our composite of GPQA Diamond, Humanity's Last Exam, AIME, and ARC-AGI-2, normalised to 0 to 100. Use it to rank, not to quote.
Context
How much text fits in one request. 1M tokens is roughly 750,000 words. Usable context is often less than the claimed figure.
$/M in
US dollars per million input tokens at list price. Output tokens cost more, usually 3 to 5 times. Open models show the cheapest hosted rate.
Capabilities
V vision input, A audio, T tool calling, R extended reasoning mode, C computer or browser use. A missing letter means not stated, not no.
AA Index
The Artificial Analysis Intelligence Index: one independently measured number across reasoning, coding, and agentic evals. The only score on this page that every model, new or old, gets on the same day.
Live rows
Found nightly on Hugging Face and OpenRouter and admitted by a Claude judge. They carry no benchmark scores until the vendor publishes them.

Video generation models

License I2V rank (1 = best)
Model Vendor License Released I2V Rank Max len (s) Resolution Audio $/gen Best for
Kling 3.0 Kuaishou closed 2026-02 1 15 (multi-shot) 1080p synced dialogue + SFX 1.00 Cinematic shots; #1 general-purpose
Veo 3.1 Google closed 2026-01 2 8 1080p yes 0.75 High-fidelity realism, prompt adherence
Sora 2 OpenAI closed 2025-10 3 12 1080p yes 0.90 Imaginative T2V; ChatGPT-integrated
Seedance 2.0 ByteDance closed 2026-04 4 15 1080p yes 0.50 Product ads, e-comm, character consistency
LTX-2 Lightricks open 2025-Q4 5 10 4K@50fps native sync 0.20 Open-weights leader; 4K + audio
Wan 2.2 Alibaba open 2025-Q3 6 6 720p N/A 0.05 Runs on a 4070; novel MoE denoiser

Image generation models

License Quality = directional leaderboard score
Model Vendor License Released Max res Quality* Speed (s) $/img Best for
Midjourney v8Midjourneyclosed2026-032K95100.04Aesthetic leader; rewritten engine, five times faster than v7
FLUX.2 [pro]Black Forest Labsclosed (API)2026-Q12K934.50.04Photoreal commercial: best skin, lighting, materials
FLUX.2 [dev]Black Forest Labsopen weights2026-Q12K8960Top open-weights photoreal model
GPT Image 2OpenAIclosed2026-Q11K+9260.04Best prompt adherence: complex composed scenes
Imagen 4 UltraGoogleclosed2025-Q42K9450.04Photoreal flagship; strong text rendering
Imagen 4 FastGoogleclosed2026-031K8620.02Cheapest fast quality at $0.02/img
Nano Banana 2Google (Gemini 3.1 Flash Image)closed2026-031K821.50.02Fastest end to end, 1 to 3 seconds; in-chat editing
Recraft V4Recraftclosed2025-Q42K8850.04Brand / design assets, vector-style outputs
Ideogram 3Ideogramclosed2025-Q41K8450.03Typography & in-image text rendering
Seedream 4.5ByteDanceclosed2026-Q12K9050.03Asian aesthetics, character consistency
Stable Diffusion 4Stability AIopen2026-Q12K8060Open ecosystem; ControlNet and LoRA backbone
Hunyuan Image 3Tencentopen2026-Q12K8350Open Tencent flagship; bilingual prompts