Generative AI Model Ranking Matrix
Living benchmark, refreshed nightly
Frontier language, video, and image generation models compared side by side. Benchmarks, capabilities, release dates, and price, so you can pick a model for the job and not the hype.
Frontier LLMs: coding and reasoning
27 benchmarked models, newest first. Then 12 newly released models tagged live, found nightly and waiting on published benchmarks. Click a column to sort. Click a model for its full profile.
Swipe sideways for benchmarks and price
| Model | Vendor | Params (B) | License | Capabilities | Released | Context | SWE-Bench Verified | SWE-Bench Pro | Terminal-Bench | Reasoning* | AA Index | $/M in | Best for |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-6 Astra | OpenAI | N/A | closed | VTRC | 2026-09-03 | 1M | N/A | N/A | N/A | N/A | N/A | 10 | OpenAI flagship. Terminal-Bench 4.0 57.7%, DeepSWE 74.1%, GPQA Diamond 96% |
| Gemini 3.8 Flash | N/A | closed | VATR | 2026-09-02 | 1M | N/A | N/A | N/A | N/A | N/A | 0.75 | Cheap agent workhorse. HLE-Verified 54.9%. Price doubles on 1 Jan 2027 | |
| Claude Fable 5.1 | Anthropic | N/A | closed | VTRC | 2026-09-01 | 1M | N/A | N/A | N/A | N/A | N/A | 10 | Top of the Claude 5 family. Terminal-Bench 4.0 55.8%, DeepSWE 67.4%, HLE with tools 65% |
| Claude Opus 5 | Anthropic | N/A | closed | VTRC | 2026-07-24 | 1M | 96 | 79.2 | N/A | N/A | N/A | 5 | SWE-Bench Verified 96.0 and Pro 79.2 at unchanged Opus pricing |
| Mistral Medium 3.5 | Mistral | 128 | open weights | TR | 2026-04-30 | 128k | 77.6 | N/A | N/A | 80 | N/A | 0.6 | 128B · SWE-V 77.6 beats Devstral 2 + Qwen3.5; τ³-Telecom 91.4 |
| Kimi K2.6 | Moonshot | 1000 | open | VTR | 2026-04-29 | 256k | 80.2 | 58.6 | 66.7 | 94 | N/A | 0.6 | 1.1T MoE · ties GPT-5.5 on SWE-Pro at $0.6/M; HLE 54 leads all |
| MiMo-V2.5-Pro | Xiaomi | 1023.2 | open | TR | 2026-04-29 | 1M | N/A | 57.2 | N/A | 89 | N/A | N/A | 1.02T MoE / 42B active · matches Opus 4.6 on SWE-Pro w/ 40% fewer tokens |
| Gemma 4 31B-it | 31 | open | VT | 2026-04-29 | 256k | 52 | N/A | N/A | 87 | N/A | N/A | 31B dense · AIME 89.2, GPQA 84.3, τ²-Retail 86.4 (12× jump from G3) | |
| Nemotron-3-Nano-Omni | NVIDIA | 30 | open | VATR | 2026-04-29 | 128k | N/A | N/A | N/A | N/A | N/A | N/A | Fresh (~17h) · 30B-A3B any-to-any reasoning |
| DeepSeek V4-Pro | DeepSeek | 861.6 | open | TR | 2026-04-27 | 128k | 80.6 | 55.4 | 67.9 | 92 | N/A | 0.9 | 1.6T MoE / 49B active · matches Opus 4.6 at fraction of cost |
| DeepSeek V4-Flash | DeepSeek | 158.1 | open | TR | 2026-04-27 | 128k | N/A | N/A | N/A | N/A | N/A | 0.3 | Fresh · 158B fast variant of V4 |
| GPT-5.5 | OpenAI | N/A | closed | VATRC | 2026-04-24 | 1M | 87.6 | 58.6 | 82.7 | 95 | N/A | 5 | SOTA Terminal-Bench 82.7; GDPval 84.9; OSWorld 78.7; ties Opus 4.7 on SWE-V |
| Qwen3.5-397B-A17B | Alibaba | 397 | open | VTR | 2026-04-24 | 256k | 80 | N/A | 54 | 92 | N/A | N/A | 403B MoE / 17B active · GPQA 88.4, MMLU-Pro 87.8, SWE-V 80.0 |
| Qwen3.6-35B-A3B | Alibaba | 35 | open | VTR | 2026-04-24 | 256k | 73.4 | 49.5 | N/A | 84 | N/A | N/A | 36B / 3B active · 73.4 SWE-V on tiny active params |
| Tencent Hy3-preview | Tencent | 298.8 | open (preview) | VTR | 2026-04-24 | 128k | 74.4 | N/A | 54.4 | 88 | N/A | N/A | 295B / 21B active · +40pts SWE-V vs Hy2; topped Tsinghua math PhD exam |
| Qwen3.6-Max-Preview | Alibaba | N/A | closed | VTR | 2026-04-20 | 1M | 79.5 | N/A | 77.1 | 89 | N/A | 6 | Agent loops: preserve_thinking across tool calls |
| MiniMax M2.7 | MiniMax | 228.7 | open | TR | 2026-04-20 | 205k | 78 | 56.22 | 57 | 87 | N/A | 0.3 | 229B · self-evolving; matches Codex on SWE-Pro |
| Claude Opus 4.7 | Anthropic | N/A | closed | VTRC | 2026-04-16 | 1M | 87.6 | 64.3 | 69.4 | 95 | N/A | 5 | Leads SWE-Bench Verified (87.6) and Pro (64.3); new tokenizer (~35% more tokens) |
| GLM-5.1 | Zhipu / Z.ai | 753.9 | open (MIT) | VTR | 2026-04-07 | 128k | 79 | 58.4 | N/A | 88 | N/A | 0.11 | 754B · #1 SWE-Bench Pro among open weights |
| Grok 4.20 Beta 2 | xAI | N/A | closed | VT | 2026-03-03 | 256k | 75 | N/A | N/A | 80 | N/A | 2 | 4-agent backbone · IFBench #1 (83); 8% of Opus cost |
| GPT-5.3 Codex | OpenAI | N/A | closed | VTRC | 2026-03 | 400k | 80 | 56.8 | 77.3 | 91 | N/A | 10 | Full SDLC agent: debug, terminal, PRDs, tests |
| Gemini 3.1 Pro | N/A | closed | VATRC | 2026-03 | 2M | 78 | N/A | N/A | 94 | N/A | 7 | Adjustable Deep Think; agentic browsing (BrowseComp 85.9) | |
| Devstral-2-123B | Mistral | 123 | open weights | TC | 2026-02-25 | 128k | N/A | N/A | N/A | N/A | N/A | 0.6 | 125B · code-specialized variant |
| Claude Opus 4.6 | Anthropic | N/A | closed | VTRC | 2026-02 | 1M | 80.8 | 57.3 | N/A | 92 | N/A | 15 | Large-codebase reasoning, multi-file refactors |
| Mistral Large 3 | Mistral | 675 | open (Apache 2.0) | VT | 2025-12-02 | 256k | N/A | N/A | N/A | 50 | N/A | 2 | 675B MoE / 41B active · non-reasoning (AIME ~40, GPQA ~44) · Apache 2.0 |
| Claude Sonnet 4.6 | Anthropic | N/A | closed | VTRC | 2025-12 | 1M | 79.6 | N/A | N/A | 86 | N/A | 3 | Default coding driver. 98% of Opus at ⅕ cost |
| Claude Haiku 4.5 | Anthropic | N/A | closed | VTRC | 2025-10-15 | 200k | 73.3 | N/A | N/A | 78 | N/A | 1 | 4 to 5 times faster than Sonnet 4.5; cheap multi-agent driver |
| Ling-3.0-flash-VL live | InclusionAI | 124 | open | VTR | 2026-09-10 | 262k | N/A | N/A | N/A | N/A | N/A | 0.00 | 124B MoE multimodal model with vision capabilities and very long context (262K), but from a less prominent lab with limited documentation; composite score of 21 meets threshold but significance is borderline. |
| Nemotron-3.5-Lightning-30B-A3B live | NVIDIA | 17.8 | open | N/A | 2026-09-10 | 1M | N/A | N/A | N/A | N/A | N/A | N/A | Novel hybrid MoE architecture (Mamba-2 + MoE + Attention) with 1M context and efficient 3B active params from a major lab, but sub-frontier scale and NVFP4 quantized variant limits composite score. |
| V4.1-Flash live | DeepSeek | 8 | open | VTR | 2026-09-10 | 1M | N/A | N/A | N/A | N/A | N/A | 0.30 | Novel Causal Encoder-Decoder (CED) architecture from a major lab is significant, but small active parameter count (8B) limits scale score. |
| gpt-6-astra-pro live | OpenAI | N/A | closed | VTR | 2026-09-04 | 1.1M | N/A | N/A | N/A | N/A | N/A | 10.00 | GPT-6 Astra Pro appears to be a frontier-class general-purpose model from OpenAI with a very large context window and high pricing indicative of a major model, served with enhanced reasoning mode. |
| GLM-5.3 live | Zhipu / Z.ai | 753.3 | open | N/A | 2026-09-04 | N/A | N/A | N/A | N/A | N/A | N/A | N/A | GLM-5.3 is a 753B-parameter open-weights MoE model achieving frontier-class coding and cyber capabilities through post-training advances, competitive with top closed models on multiple benchmarks. |
| qwen3.8-max-0902 live | Alibaba | N/A | closed | VTR | 2026-09-03 | 1M | N/A | N/A | N/A | N/A | N/A | 2.00 | 2.4T parameter MoE model with multimodal capabilities (text/image/video) and 1M context window represents frontier-class scale and capability. |
| Hy4-preview live | Tencent | 780 | open | N/A | 2026-08-28 | N/A | N/A | N/A | N/A | N/A | N/A | N/A | Hy4-preview is a 780B MoE flagship model from Tencent, representing frontier-class scale and a new generation of their Hunyuan series with MoE architecture. |
| GLM-5.3-Flash live | Zhipu / Z.ai | N/A | open | VTR | 2026-08-26 | 1.3M | N/A | N/A | N/A | N/A | N/A | 0.15 | GLM-5.3-Flash features a novel hybrid sparse/linear attention architecture with a massive 1.3M token context and native multimodal support, but parameter count is unknown and it appears to be a flash/efficient variant rather than a frontier-scale model. |
| V4-Flash-Vision-Exp live | DeepSeek | N/A | open | VTR | 2026-08-21 | 1M | N/A | N/A | N/A | N/A | N/A | 0.22 | DeepSeek V4 Flash Vision Exp is a frontier-class multimodal model with a massive 1M token context, but parameter count is unknown and it is experimental/vision-extended variant rather than a fully novel architecture. |
| glm-latest live | Zhipu / Z.ai | N/A | closed | TR | 2026-08-19 | 1.3M | N/A | N/A | N/A | N/A | N/A | 1.00 | GLM-latest from Z.ai is a frontier-class general-purpose model with an exceptionally long context window (1.3M tokens), but unknown parameter count, proprietary/closed nature, and minimal documentation make full assessment difficult. |
| Qwen3.8-27B live | Alibaba | 27 | open | VTR | 2026-08-14 | 1M | N/A | N/A | N/A | N/A | N/A | 0.42 | 27B dense vision-language model with 1M context and flexible thinking is solid but below frontier-class composite threshold for automatic inclusion. |
| V4-Flash-0731 live | DeepSeek | 304.2 | open | N/A | 2026-08-01 | N/A | N/A | N/A | N/A | N/A | N/A | N/A | DeepSeek-V4-Flash-0731 is a frontier-class MoE model from a major lab with 304B total parameters, high download counts, and an official technical report, representing a significant general-purpose release. |
OpenAI flagship. Terminal-Bench 4.0 57.7%, DeepSWE 74.1%, GPQA Diamond 96%. Text, image, and file input with tool calling, extended reasoning, and computer use. OpenAI reports OSWorld 2.0 at 72.6% and GPQA Diamond at 96.0%.
- SWE-Bench Verified
- N/A
- SWE-Bench Pro
- N/A
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 1M
- Price in / out
- $10 / $50 per M
Cheap agent workhorse. HLE-Verified 54.9%. Price doubles on 1 Jan 2027. Text, image, audio, video, and PDF input with tool calling and extended reasoning. Google reports HLE-Verified at 54.9%.
- SWE-Bench Verified
- N/A
- SWE-Bench Pro
- N/A
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 1M
- Price in / out
- $0.75 / $3.75 per M
Top of the Claude 5 family. Terminal-Bench 4.0 55.8%, DeepSWE 67.4%, HLE with tools 65%. Text, image, and file input with tool calling, extended reasoning, and computer use. Built for multi-hour agentic sessions. Cache reads cut to $0.25 per million.
- SWE-Bench Verified
- N/A
- SWE-Bench Pro
- N/A
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 1M
- Price in / out
- $10 / $50 per M
SWE-Bench Verified 96.0 and Pro 79.2 at unchanged Opus pricing. Text and image input with tool calling, extended thinking on by default, and computer use. Anthropic reports OSWorld 2.0 at 70.6%.
- SWE-Bench Verified
- 96%
- SWE-Bench Pro
- 79.2%
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 1M
- Price in / out
- $5 / $25 per M
128B · SWE-V 77.6 beats Devstral 2 + Qwen3.5; τ³-Telecom 91.4. Mistral Medium 3.5 supports tool calling and extended reasoning, with a 128k-token context window.
- SWE-Bench Verified
- 77.6%
- SWE-Bench Pro
- N/A
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 128k
- Price in / out
- $0.6 / N/A per M
1.1T MoE · ties GPT-5.5 on SWE-Pro at $0.6/M; HLE 54 leads all. Kimi K2.6 supports vision input, tool calling and extended reasoning, with a 256k-token context window.
- SWE-Bench Verified
- 80.2%
- SWE-Bench Pro
- 58.6%
- Terminal-Bench
- 66.7%
- AA Index
- N/A
- Context
- 256k
- Price in / out
- $0.6 / N/A per M
1.02T MoE / 42B active · matches Opus 4.6 on SWE-Pro w/ 40% fewer tokens. MiMo-V2.5-Pro supports tool calling and extended reasoning, with a 1M-token context window.
- SWE-Bench Verified
- N/A
- SWE-Bench Pro
- 57.2%
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 1M
- Price in / out
- N/A / N/A per M
31B dense · AIME 89.2, GPQA 84.3, τ²-Retail 86.4 (12× jump from G3). Gemma 4 31B-it supports vision input and tool calling, with a 256k-token context window.
- SWE-Bench Verified
- 52%
- SWE-Bench Pro
- N/A
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 256k
- Price in / out
- N/A / N/A per M
Fresh (~17h) · 30B-A3B any-to-any reasoning. Nemotron-3-Nano-Omni supports vision input, audio, tool calling and extended reasoning, with a 128k-token context window.
- SWE-Bench Verified
- N/A
- SWE-Bench Pro
- N/A
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 128k
- Price in / out
- N/A / N/A per M
1.6T MoE / 49B active · matches Opus 4.6 at fraction of cost. DeepSeek V4-Pro supports tool calling and extended reasoning, with a 128k-token context window.
- SWE-Bench Verified
- 80.6%
- SWE-Bench Pro
- 55.4%
- Terminal-Bench
- 67.9%
- AA Index
- N/A
- Context
- 128k
- Price in / out
- $0.9 / N/A per M
Fresh · 158B fast variant of V4. DeepSeek V4-Flash supports tool calling and extended reasoning, with a 128k-token context window.
- SWE-Bench Verified
- N/A
- SWE-Bench Pro
- N/A
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 128k
- Price in / out
- $0.3 / N/A per M
SOTA Terminal-Bench 82.7; GDPval 84.9; OSWorld 78.7; ties Opus 4.7 on SWE-V. GPT-5.5 supports vision input, audio, tool calling, extended reasoning and computer use, with a 1M-token context window.
- SWE-Bench Verified
- 87.6%
- SWE-Bench Pro
- 58.6%
- Terminal-Bench
- 82.7%
- AA Index
- N/A
- Context
- 1M
- Price in / out
- $5 / $30 per M
403B MoE / 17B active · GPQA 88.4, MMLU-Pro 87.8, SWE-V 80.0. Qwen3.5-397B-A17B supports vision input, tool calling and extended reasoning, with a 256k-token context window.
- SWE-Bench Verified
- 80%
- SWE-Bench Pro
- N/A
- Terminal-Bench
- 54%
- AA Index
- N/A
- Context
- 256k
- Price in / out
- N/A / N/A per M
36B / 3B active · 73.4 SWE-V on tiny active params. Qwen3.6-35B-A3B supports vision input, tool calling and extended reasoning, with a 256k-token context window.
- SWE-Bench Verified
- 73.4%
- SWE-Bench Pro
- 49.5%
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 256k
- Price in / out
- N/A / N/A per M
295B / 21B active · +40pts SWE-V vs Hy2; topped Tsinghua math PhD exam. Tencent Hy3-preview supports vision input, tool calling and extended reasoning, with a 128k-token context window.
- SWE-Bench Verified
- 74.4%
- SWE-Bench Pro
- N/A
- Terminal-Bench
- 54.4%
- AA Index
- N/A
- Context
- 128k
- Price in / out
- N/A / N/A per M
Agent loops: preserve_thinking across tool calls. Qwen3.6-Max-Preview supports vision input, tool calling and extended reasoning, with a 1M-token context window.
- SWE-Bench Verified
- 79.5%
- SWE-Bench Pro
- N/A
- Terminal-Bench
- 77.1%
- AA Index
- N/A
- Context
- 1M
- Price in / out
- $6 / N/A per M
229B · self-evolving; matches Codex on SWE-Pro. MiniMax M2.7 supports tool calling and extended reasoning, with a 205k-token context window.
- SWE-Bench Verified
- 78%
- SWE-Bench Pro
- 56.22%
- Terminal-Bench
- 57%
- AA Index
- N/A
- Context
- 205k
- Price in / out
- $0.3 / N/A per M
Leads SWE-Bench Verified (87.6) and Pro (64.3); new tokenizer (~35% more tokens). Claude Opus 4.7 supports vision input, tool calling, extended reasoning and computer use, with a 1M-token context window.
- SWE-Bench Verified
- 87.6%
- SWE-Bench Pro
- 64.3%
- Terminal-Bench
- 69.4%
- AA Index
- N/A
- Context
- 1M
- Price in / out
- $5 / $25 per M
754B · #1 SWE-Bench Pro among open weights. GLM-5.1 supports vision input, tool calling and extended reasoning, with a 128k-token context window.
- SWE-Bench Verified
- 79%
- SWE-Bench Pro
- 58.4%
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 128k
- Price in / out
- $0.11 / N/A per M
4-agent backbone · IFBench #1 (83); 8% of Opus cost. Grok 4.20 Beta 2 supports vision input and tool calling, with a 256k-token context window.
- SWE-Bench Verified
- 75%
- SWE-Bench Pro
- N/A
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 256k
- Price in / out
- $2 / N/A per M
Full SDLC agent: debug, terminal, PRDs, tests. GPT-5.3 Codex supports vision input, tool calling, extended reasoning and computer use, with a 400k-token context window.
- SWE-Bench Verified
- 80%
- SWE-Bench Pro
- 56.8%
- Terminal-Bench
- 77.3%
- AA Index
- N/A
- Context
- 400k
- Price in / out
- $10 / N/A per M
Adjustable Deep Think; agentic browsing (BrowseComp 85.9). Gemini 3.1 Pro supports vision input, audio, tool calling, extended reasoning and computer use, with a 2M-token context window.
- SWE-Bench Verified
- 78%
- SWE-Bench Pro
- N/A
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 2M
- Price in / out
- $7 / N/A per M
125B · code-specialized variant. Devstral-2-123B supports tool calling and computer use, with a 128k-token context window.
- SWE-Bench Verified
- N/A
- SWE-Bench Pro
- N/A
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 128k
- Price in / out
- $0.6 / N/A per M
Large-codebase reasoning, multi-file refactors. Claude Opus 4.6 supports vision input, tool calling, extended reasoning and computer use, with a 1M-token context window.
- SWE-Bench Verified
- 80.8%
- SWE-Bench Pro
- 57.3%
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 1M
- Price in / out
- $15 / N/A per M
675B MoE / 41B active · non-reasoning (AIME ~40, GPQA ~44) · Apache 2.0. Mistral Large 3 supports vision input and tool calling, with a 256k-token context window.
- SWE-Bench Verified
- N/A
- SWE-Bench Pro
- N/A
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 256k
- Price in / out
- $2 / N/A per M
Default coding driver. 98% of Opus at ⅕ cost. Claude Sonnet 4.6 supports vision input, tool calling, extended reasoning and computer use, with a 1M-token context window.
- SWE-Bench Verified
- 79.6%
- SWE-Bench Pro
- N/A
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 1M
- Price in / out
- $3 / $15 per M
4 to 5 times faster than Sonnet 4.5; cheap multi-agent driver. Claude Haiku 4.5 supports vision input, tool calling, extended reasoning and computer use, with a 200k-token context window.
- SWE-Bench Verified
- 73.3%
- SWE-Bench Pro
- N/A
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 200k
- Price in / out
- $1 / $5 per M
124B MoE multimodal model with vision capabilities and very long context (262K), but from a less prominent lab with limited documentation; composite score of 21 meets threshold but significance is borderline.
- AA Index
- Awaiting benchmarks
- Context
- 262k
- Price in / out
- $0 / $0 per M
- Found via
- OpenRouter
Nemotron-3.5-Lightning-30B-A3B
Novel hybrid MoE architecture (Mamba-2 + MoE + Attention) with 1M context and efficient 3B active params from a major lab, but sub-frontier scale and NVFP4 quantized variant limits composite score.
- AA Index
- Awaiting benchmarks
- Context
- 1M
- Price in / out
- N/A / N/A per M
- Found via
- Hugging Face
Novel Causal Encoder-Decoder (CED) architecture from a major lab is significant, but small active parameter count (8B) limits scale score.
- AA Index
- Awaiting benchmarks
- Context
- 1M
- Price in / out
- $0.3 / $1.2 per M
- Found via
- OpenRouter
GPT-6 Astra Pro appears to be a frontier-class general-purpose model from OpenAI with a very large context window and high pricing indicative of a major model, served with enhanced reasoning mode.
- AA Index
- Awaiting benchmarks
- Context
- 1.1M
- Price in / out
- $10 / $50 per M
- Found via
- OpenRouter
GLM-5.3 is a 753B-parameter open-weights MoE model achieving frontier-class coding and cyber capabilities through post-training advances, competitive with top closed models on multiple benchmarks.
- AA Index
- Awaiting benchmarks
- Context
- N/A
- Price in / out
- N/A / N/A per M
- Found via
- Hugging Face
2.4T parameter MoE model with multimodal capabilities (text/image/video) and 1M context window represents frontier-class scale and capability.
- AA Index
- Awaiting benchmarks
- Context
- 1M
- Price in / out
- $2 / $6 per M
- Found via
- OpenRouter
Hy4-preview is a 780B MoE flagship model from Tencent, representing frontier-class scale and a new generation of their Hunyuan series with MoE architecture.
- AA Index
- Awaiting benchmarks
- Context
- N/A
- Price in / out
- N/A / N/A per M
- Found via
- Hugging Face
GLM-5.3-Flash features a novel hybrid sparse/linear attention architecture with a massive 1.3M token context and native multimodal support, but parameter count is unknown and it appears to be a flash/efficient variant rather than a frontier-scale model.
- AA Index
- Awaiting benchmarks
- Context
- 1.3M
- Price in / out
- $0.15 / $0.5 per M
- Found via
- OpenRouter
DeepSeek V4 Flash Vision Exp is a frontier-class multimodal model with a massive 1M token context, but parameter count is unknown and it is experimental/vision-extended variant rather than a fully novel architecture.
- AA Index
- Awaiting benchmarks
- Context
- 1M
- Price in / out
- $0.22 / $0.66 per M
- Found via
- OpenRouter
GLM-latest from Z.ai is a frontier-class general-purpose model with an exceptionally long context window (1.3M tokens), but unknown parameter count, proprietary/closed nature, and minimal documentation make full assessment difficult.
- AA Index
- Awaiting benchmarks
- Context
- 1.3M
- Price in / out
- $1 / $3.41 per M
- Found via
- OpenRouter
27B dense vision-language model with 1M context and flexible thinking is solid but below frontier-class composite threshold for automatic inclusion.
- AA Index
- Awaiting benchmarks
- Context
- 1M
- Price in / out
- $0.42 / $3 per M
- Found via
- OpenRouter
DeepSeek-V4-Flash-0731 is a frontier-class MoE model from a major lab with 304B total parameters, high download counts, and an official technical report, representing a significant general-purpose release.
- AA Index
- Awaiting benchmarks
- Context
- N/A
- Price in / out
- N/A / N/A per M
- Found via
- Hugging Face
What the columns mean
- SWE-Bench Verified
- Share of 500 real GitHub issues the model fixed end to end, tests passing. The best single signal for coding agents.
- SWE-Bench Pro
- A harder, contamination-resistant set of long-horizon coding tasks. Scores run lower than Verified for every model.
- Terminal-Bench
- Tasks completed in a real shell: builds, data wrangling, debugging. Measures agentic work, not chat.
- Reasoning
- Our composite of GPQA Diamond, Humanity's Last Exam, AIME, and ARC-AGI-2, normalised to 0 to 100. Use it to rank, not to quote.
- Context
- How much text fits in one request. 1M tokens is roughly 750,000 words. Usable context is often less than the claimed figure.
- $/M in
- US dollars per million input tokens at list price. Output tokens cost more, usually 3 to 5 times. Open models show the cheapest hosted rate.
- Capabilities
- V vision input, A audio, T tool calling, R extended reasoning mode, C computer or browser use. A missing letter means not stated, not no.
- AA Index
- The Artificial Analysis Intelligence Index: one independently measured number across reasoning, coding, and agentic evals. The only score on this page that every model, new or old, gets on the same day.
- Live rows
- Found nightly on Hugging Face and OpenRouter and admitted by a Claude judge. They carry no benchmark scores until the vendor publishes them.
Video generation models
| Model | Vendor | License | Released | I2V Rank | Max len (s) | Resolution | Audio | $/gen | Best for |
|---|---|---|---|---|---|---|---|---|---|
| Kling 3.0 | Kuaishou | closed | 2026-02 | 1 | 15 (multi-shot) | 1080p | synced dialogue + SFX | 1.00 | Cinematic shots; #1 general-purpose |
| Veo 3.1 | closed | 2026-01 | 2 | 8 | 1080p | yes | 0.75 | High-fidelity realism, prompt adherence | |
| Sora 2 | OpenAI | closed | 2025-10 | 3 | 12 | 1080p | yes | 0.90 | Imaginative T2V; ChatGPT-integrated |
| Seedance 2.0 | ByteDance | closed | 2026-04 | 4 | 15 | 1080p | yes | 0.50 | Product ads, e-comm, character consistency |
| LTX-2 | Lightricks | open | 2025-Q4 | 5 | 10 | 4K@50fps | native sync | 0.20 | Open-weights leader; 4K + audio |
| Wan 2.2 | Alibaba | open | 2025-Q3 | 6 | 6 | 720p | N/A | 0.05 | Runs on a 4070; novel MoE denoiser |
Image generation models
| Model | Vendor | License | Released | Max res | Quality* | Speed (s) | $/img | Best for |
|---|---|---|---|---|---|---|---|---|
| Midjourney v8 | Midjourney | closed | 2026-03 | 2K | 95 | 10 | 0.04 | Aesthetic leader; rewritten engine, five times faster than v7 |
| FLUX.2 [pro] | Black Forest Labs | closed (API) | 2026-Q1 | 2K | 93 | 4.5 | 0.04 | Photoreal commercial: best skin, lighting, materials |
| FLUX.2 [dev] | Black Forest Labs | open weights | 2026-Q1 | 2K | 89 | 6 | 0 | Top open-weights photoreal model |
| GPT Image 2 | OpenAI | closed | 2026-Q1 | 1K+ | 92 | 6 | 0.04 | Best prompt adherence: complex composed scenes |
| Imagen 4 Ultra | closed | 2025-Q4 | 2K | 94 | 5 | 0.04 | Photoreal flagship; strong text rendering | |
| Imagen 4 Fast | closed | 2026-03 | 1K | 86 | 2 | 0.02 | Cheapest fast quality at $0.02/img | |
| Nano Banana 2 | Google (Gemini 3.1 Flash Image) | closed | 2026-03 | 1K | 82 | 1.5 | 0.02 | Fastest end to end, 1 to 3 seconds; in-chat editing |
| Recraft V4 | Recraft | closed | 2025-Q4 | 2K | 88 | 5 | 0.04 | Brand / design assets, vector-style outputs |
| Ideogram 3 | Ideogram | closed | 2025-Q4 | 1K | 84 | 5 | 0.03 | Typography & in-image text rendering |
| Seedream 4.5 | ByteDance | closed | 2026-Q1 | 2K | 90 | 5 | 0.03 | Asian aesthetics, character consistency |
| Stable Diffusion 4 | Stability AI | open | 2026-Q1 | 2K | 80 | 6 | 0 | Open ecosystem; ControlNet and LoRA backbone |
| Hunyuan Image 3 | Tencent | open | 2026-Q1 | 2K | 83 | 5 | 0 | Open Tencent flagship; bilingual prompts |