Generative AI Model Ranking Matrix
Living benchmark, refreshed nightly
Frontier language, video, and image generation models compared side by side. Benchmarks, capabilities, release dates, and price, so you can pick a model for the job and not the hype.
Frontier LLMs: coding and reasoning
27 benchmarked models, newest first. Then 13 newly released models tagged live, found nightly and waiting on published benchmarks. Click a column to sort. Click a model for its full profile.
Swipe sideways for benchmarks and price
| Model | Vendor | Params (B) | License | Capabilities | Released | Context | SWE-Bench Verified | SWE-Bench Pro | Terminal-Bench | Reasoning* | AA Index | $/M in | Best for |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-6 Astra | OpenAI | N/A | closed | VTRC | 2026-09-03 | 1M | N/A | N/A | N/A | N/A | N/A | 10 | OpenAI flagship. Terminal-Bench 4.0 57.7%, DeepSWE 74.1%, GPQA Diamond 96% |
| Gemini 3.8 Flash | N/A | closed | VATR | 2026-09-02 | 1M | N/A | N/A | N/A | N/A | N/A | 0.75 | Cheap agent workhorse. HLE-Verified 54.9%. Price doubles on 1 Jan 2027 | |
| Claude Fable 5.1 | Anthropic | N/A | closed | VTRC | 2026-09-01 | 1M | N/A | N/A | N/A | N/A | N/A | 10 | Top of the Claude 5 family. Terminal-Bench 4.0 55.8%, DeepSWE 67.4%, HLE with tools 65% |
| Claude Opus 5 | Anthropic | N/A | closed | VTRC | 2026-07-24 | 1M | 96 | 79.2 | N/A | N/A | N/A | 5 | SWE-Bench Verified 96.0 and Pro 79.2 at unchanged Opus pricing |
| Mistral Medium 3.5 | Mistral | 128 | open weights | TR | 2026-04-30 | 128k | 77.6 | N/A | N/A | 80 | N/A | 0.6 | 128B · SWE-V 77.6 beats Devstral 2 + Qwen3.5; τ³-Telecom 91.4 |
| Kimi K2.6 | Moonshot | 1000 | open | VTR | 2026-04-29 | 256k | 80.2 | 58.6 | 66.7 | 94 | N/A | 0.6 | 1.1T MoE · ties GPT-5.5 on SWE-Pro at $0.6/M; HLE 54 leads all |
| MiMo-V2.5-Pro | Xiaomi | 1023.2 | open | TR | 2026-04-29 | 1M | N/A | 57.2 | N/A | 89 | N/A | N/A | 1.02T MoE / 42B active · matches Opus 4.6 on SWE-Pro w/ 40% fewer tokens |
| Gemma 4 31B-it | 31 | open | VT | 2026-04-29 | 256k | 52 | N/A | N/A | 87 | N/A | N/A | 31B dense · AIME 89.2, GPQA 84.3, τ²-Retail 86.4 (12× jump from G3) | |
| Nemotron-3-Nano-Omni | NVIDIA | 30 | open | VATR | 2026-04-29 | 128k | N/A | N/A | N/A | N/A | N/A | N/A | Fresh (~17h) · 30B-A3B any-to-any reasoning |
| DeepSeek V4-Pro | DeepSeek | 861.6 | open | TR | 2026-04-27 | 128k | 80.6 | 55.4 | 67.9 | 92 | N/A | 0.9 | 1.6T MoE / 49B active · matches Opus 4.6 at fraction of cost |
| DeepSeek V4-Flash | DeepSeek | 158.1 | open | TR | 2026-04-27 | 128k | N/A | N/A | N/A | N/A | N/A | 0.3 | Fresh · 158B fast variant of V4 |
| GPT-5.5 | OpenAI | N/A | closed | VATRC | 2026-04-24 | 1M | 87.6 | 58.6 | 82.7 | 95 | N/A | 5 | SOTA Terminal-Bench 82.7; GDPval 84.9; OSWorld 78.7; ties Opus 4.7 on SWE-V |
| Qwen3.5-397B-A17B | Alibaba | 397 | open | VTR | 2026-04-24 | 256k | 80 | N/A | 54 | 92 | N/A | N/A | 403B MoE / 17B active · GPQA 88.4, MMLU-Pro 87.8, SWE-V 80.0 |
| Qwen3.6-35B-A3B | Alibaba | 35 | open | VTR | 2026-04-24 | 256k | 73.4 | 49.5 | N/A | 84 | N/A | N/A | 36B / 3B active · 73.4 SWE-V on tiny active params |
| Tencent Hy3-preview | Tencent | 298.8 | open (preview) | VTR | 2026-04-24 | 128k | 74.4 | N/A | 54.4 | 88 | N/A | N/A | 295B / 21B active · +40pts SWE-V vs Hy2; topped Tsinghua math PhD exam |
| Qwen3.6-Max-Preview | Alibaba | N/A | closed | VTR | 2026-04-20 | 1M | 79.5 | N/A | 77.1 | 89 | N/A | 6 | Agent loops: preserve_thinking across tool calls |
| MiniMax M2.7 | MiniMax | 228.7 | open | TR | 2026-04-20 | 205k | 78 | 56.22 | 57 | 87 | N/A | 0.3 | 229B · self-evolving; matches Codex on SWE-Pro |
| Claude Opus 4.7 | Anthropic | N/A | closed | VTRC | 2026-04-16 | 1M | 87.6 | 64.3 | 69.4 | 95 | N/A | 5 | Leads SWE-Bench Verified (87.6) and Pro (64.3); new tokenizer (~35% more tokens) |
| GLM-5.1 | Zhipu / Z.ai | 753.9 | open (MIT) | VTR | 2026-04-07 | 128k | 79 | 58.4 | N/A | 88 | N/A | 0.11 | 754B · #1 SWE-Bench Pro among open weights |
| Grok 4.20 Beta 2 | xAI | N/A | closed | VT | 2026-03-03 | 256k | 75 | N/A | N/A | 80 | N/A | 2 | 4-agent backbone · IFBench #1 (83); 8% of Opus cost |
| GPT-5.3 Codex | OpenAI | N/A | closed | VTRC | 2026-03 | 400k | 80 | 56.8 | 77.3 | 91 | N/A | 10 | Full SDLC agent: debug, terminal, PRDs, tests |
| Gemini 3.1 Pro | N/A | closed | VATRC | 2026-03 | 2M | 78 | N/A | N/A | 94 | N/A | 7 | Adjustable Deep Think; agentic browsing (BrowseComp 85.9) | |
| Devstral-2-123B | Mistral | 123 | open weights | TC | 2026-02-25 | 128k | N/A | N/A | N/A | N/A | N/A | 0.6 | 125B · code-specialized variant |
| Claude Opus 4.6 | Anthropic | N/A | closed | VTRC | 2026-02 | 1M | 80.8 | 57.3 | N/A | 92 | N/A | 15 | Large-codebase reasoning, multi-file refactors |
| Mistral Large 3 | Mistral | 675 | open (Apache 2.0) | VT | 2025-12-02 | 256k | N/A | N/A | N/A | 50 | N/A | 2 | 675B MoE / 41B active · non-reasoning (AIME ~40, GPQA ~44) · Apache 2.0 |
| Claude Sonnet 4.6 | Anthropic | N/A | closed | VTRC | 2025-12 | 1M | 79.6 | N/A | N/A | 86 | N/A | 3 | Default coding driver. 98% of Opus at ⅕ cost |
| Claude Haiku 4.5 | Anthropic | N/A | closed | VTRC | 2025-10-15 | 200k | 73.3 | N/A | N/A | 78 | N/A | 1 | 4 to 5 times faster than Sonnet 4.5; cheap multi-agent driver |
| gpt-astra-latest live | OpenAI | N/A | closed | VTR | 2026-09-11 | 1.1M | N/A | N/A | N/A | N/A | N/A | 10.00 | Insufficient information to confirm frontier-class capabilities; model card provides no substantive details beyond routing to a model family, making full evaluation impossible. |
| gpt-sol-latest live | OpenAI | N/A | closed | VTR | 2026-09-11 | 1.1M | N/A | N/A | N/A | N/A | N/A | 2.00 | Insufficient information to assess the model's true capabilities, architecture, or scale; the model card only states it redirects to the latest GPT Sol family model with no further details. |
| gpt-terra-latest live | OpenAI | N/A | closed | VTR | 2026-09-11 | 1.1M | N/A | N/A | N/A | N/A | N/A | 2.00 | Insufficient information to assess the model's true capabilities, architecture, or scale; the model card provides only a routing description with no technical details. |
| gpt-luna-latest live | OpenAI | N/A | closed | VTR | 2026-09-11 | 1.1M | N/A | N/A | N/A | N/A | N/A | 0.20 | Insufficient information to determine if this is a frontier-class model; the model card provides almost no technical details, params are unknown, and 'GPT Luna' is not a recognized public OpenAI release as of the knowledge cutoff. |
| Nex-N2.5-Pro live | nex-agi | 396.8 | open | VT | 2026-09-11 | N/A | N/A | N/A | N/A | N/A | N/A | N/A | Nex-N2.5-Pro is a ~397B MoE agentic model with multimodal (image+text) capabilities, computer use, and web browsing, representing a meaningful frontier-class post-training effort on a large-scale foundation. |
| V4.1-Flash live | DeepSeek | 8 | open | VTR | 2026-09-10 | 1M | N/A | N/A | N/A | N/A | N/A | 0.15 | Novel Causal Encoder-Decoder (CED) architecture from a major lab is significant, but small active parameter count (8B) limits scale score despite the architectural innovation. |
| gpt-6-astra-pro live | OpenAI | N/A | closed | VTR | 2026-09-04 | 1.1M | N/A | N/A | N/A | N/A | N/A | 10.00 | GPT-6 Astra Pro appears to be a frontier-class model from OpenAI with a very large context window and high pricing suggesting significant scale, but it is merely a reasoning-mode variant of GPT-6 Astra rather than a distinct model, and no verifiable metadata (params, architecture, capabilities) is available to confirm claims. |
| GLM-5.3 live | Zhipu / Z.ai | 753.3 | open | N/A | 2026-09-04 | N/A | N/A | N/A | N/A | N/A | N/A | N/A | GLM-5.3 is a 753B-parameter open-weights MoE model achieving frontier-class coding and cyber capability benchmarks, representing a significant post-training advancement over GLM-5.2 with SOTA results on multiple public benchmarks. |
| qwen3.8-max-0902 live | Alibaba | N/A | closed | VTR | 2026-09-03 | 1M | N/A | N/A | N/A | N/A | N/A | 2.00 | 2.4T parameter MoE model with multimodal capabilities (text/image/video) and 1M context window represents frontier-class scale and capability. |
| Hy4-preview live | Tencent | 780 | open | N/A | 2026-08-28 | N/A | N/A | N/A | N/A | N/A | N/A | N/A | Tencent's 780B MoE flagship model represents frontier-class scale and capability with a new-generation architecture from a major lab. |
| Qwen3.8-Flash-Next live | Alibaba | N/A | open | VTR | 2026-08-26 | 1M | N/A | N/A | N/A | N/A | N/A | 0.15 | Qwen3.8 Flash appears to be a capable multimodal reasoning model with a 1M context window and broad capabilities, but parameter count is unknown and it may be a smaller/flash variant, keeping composite below include threshold. |
| GLM-5.3-Flash live | Zhipu / Z.ai | N/A | open | VTR | 2026-08-26 | 1.3M | N/A | N/A | N/A | N/A | N/A | 0.15 | GLM-5.3-Flash features a novel hybrid sparse/linear attention architecture with a very large 1.3M token context window suited for agent tasks, but parameter count is unknown and it's a flash/efficient variant rather than a frontier-scale model. |
| V4-Flash-0731 live | DeepSeek | 304.2 | open | N/A | 2026-08-01 | N/A | N/A | N/A | N/A | N/A | N/A | N/A | DeepSeek-V4-Flash-0731 is a frontier-class 304B-parameter MoE model from a major lab with a dedicated technical report, high download counts, and represents a new official release superseding a preview version. |
OpenAI flagship. Terminal-Bench 4.0 57.7%, DeepSWE 74.1%, GPQA Diamond 96%. Text, image, and file input with tool calling, extended reasoning, and computer use. OpenAI reports OSWorld 2.0 at 72.6% and GPQA Diamond at 96.0%.
- SWE-Bench Verified
- N/A
- SWE-Bench Pro
- N/A
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 1M
- Price in / out
- $10 / $50 per M
Cheap agent workhorse. HLE-Verified 54.9%. Price doubles on 1 Jan 2027. Text, image, audio, video, and PDF input with tool calling and extended reasoning. Google reports HLE-Verified at 54.9%.
- SWE-Bench Verified
- N/A
- SWE-Bench Pro
- N/A
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 1M
- Price in / out
- $0.75 / $3.75 per M
Top of the Claude 5 family. Terminal-Bench 4.0 55.8%, DeepSWE 67.4%, HLE with tools 65%. Text, image, and file input with tool calling, extended reasoning, and computer use. Built for multi-hour agentic sessions. Cache reads cut to $0.25 per million.
- SWE-Bench Verified
- N/A
- SWE-Bench Pro
- N/A
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 1M
- Price in / out
- $10 / $50 per M
SWE-Bench Verified 96.0 and Pro 79.2 at unchanged Opus pricing. Text and image input with tool calling, extended thinking on by default, and computer use. Anthropic reports OSWorld 2.0 at 70.6%.
- SWE-Bench Verified
- 96%
- SWE-Bench Pro
- 79.2%
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 1M
- Price in / out
- $5 / $25 per M
128B · SWE-V 77.6 beats Devstral 2 + Qwen3.5; τ³-Telecom 91.4. Mistral Medium 3.5 supports tool calling and extended reasoning, with a 128k-token context window.
- SWE-Bench Verified
- 77.6%
- SWE-Bench Pro
- N/A
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 128k
- Price in / out
- $0.6 / N/A per M
1.1T MoE · ties GPT-5.5 on SWE-Pro at $0.6/M; HLE 54 leads all. Kimi K2.6 supports vision input, tool calling and extended reasoning, with a 256k-token context window.
- SWE-Bench Verified
- 80.2%
- SWE-Bench Pro
- 58.6%
- Terminal-Bench
- 66.7%
- AA Index
- N/A
- Context
- 256k
- Price in / out
- $0.6 / N/A per M
1.02T MoE / 42B active · matches Opus 4.6 on SWE-Pro w/ 40% fewer tokens. MiMo-V2.5-Pro supports tool calling and extended reasoning, with a 1M-token context window.
- SWE-Bench Verified
- N/A
- SWE-Bench Pro
- 57.2%
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 1M
- Price in / out
- N/A / N/A per M
31B dense · AIME 89.2, GPQA 84.3, τ²-Retail 86.4 (12× jump from G3). Gemma 4 31B-it supports vision input and tool calling, with a 256k-token context window.
- SWE-Bench Verified
- 52%
- SWE-Bench Pro
- N/A
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 256k
- Price in / out
- N/A / N/A per M
Fresh (~17h) · 30B-A3B any-to-any reasoning. Nemotron-3-Nano-Omni supports vision input, audio, tool calling and extended reasoning, with a 128k-token context window.
- SWE-Bench Verified
- N/A
- SWE-Bench Pro
- N/A
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 128k
- Price in / out
- N/A / N/A per M
1.6T MoE / 49B active · matches Opus 4.6 at fraction of cost. DeepSeek V4-Pro supports tool calling and extended reasoning, with a 128k-token context window.
- SWE-Bench Verified
- 80.6%
- SWE-Bench Pro
- 55.4%
- Terminal-Bench
- 67.9%
- AA Index
- N/A
- Context
- 128k
- Price in / out
- $0.9 / N/A per M
Fresh · 158B fast variant of V4. DeepSeek V4-Flash supports tool calling and extended reasoning, with a 128k-token context window.
- SWE-Bench Verified
- N/A
- SWE-Bench Pro
- N/A
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 128k
- Price in / out
- $0.3 / N/A per M
SOTA Terminal-Bench 82.7; GDPval 84.9; OSWorld 78.7; ties Opus 4.7 on SWE-V. GPT-5.5 supports vision input, audio, tool calling, extended reasoning and computer use, with a 1M-token context window.
- SWE-Bench Verified
- 87.6%
- SWE-Bench Pro
- 58.6%
- Terminal-Bench
- 82.7%
- AA Index
- N/A
- Context
- 1M
- Price in / out
- $5 / $30 per M
403B MoE / 17B active · GPQA 88.4, MMLU-Pro 87.8, SWE-V 80.0. Qwen3.5-397B-A17B supports vision input, tool calling and extended reasoning, with a 256k-token context window.
- SWE-Bench Verified
- 80%
- SWE-Bench Pro
- N/A
- Terminal-Bench
- 54%
- AA Index
- N/A
- Context
- 256k
- Price in / out
- N/A / N/A per M
36B / 3B active · 73.4 SWE-V on tiny active params. Qwen3.6-35B-A3B supports vision input, tool calling and extended reasoning, with a 256k-token context window.
- SWE-Bench Verified
- 73.4%
- SWE-Bench Pro
- 49.5%
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 256k
- Price in / out
- N/A / N/A per M
295B / 21B active · +40pts SWE-V vs Hy2; topped Tsinghua math PhD exam. Tencent Hy3-preview supports vision input, tool calling and extended reasoning, with a 128k-token context window.
- SWE-Bench Verified
- 74.4%
- SWE-Bench Pro
- N/A
- Terminal-Bench
- 54.4%
- AA Index
- N/A
- Context
- 128k
- Price in / out
- N/A / N/A per M
Agent loops: preserve_thinking across tool calls. Qwen3.6-Max-Preview supports vision input, tool calling and extended reasoning, with a 1M-token context window.
- SWE-Bench Verified
- 79.5%
- SWE-Bench Pro
- N/A
- Terminal-Bench
- 77.1%
- AA Index
- N/A
- Context
- 1M
- Price in / out
- $6 / N/A per M
229B · self-evolving; matches Codex on SWE-Pro. MiniMax M2.7 supports tool calling and extended reasoning, with a 205k-token context window.
- SWE-Bench Verified
- 78%
- SWE-Bench Pro
- 56.22%
- Terminal-Bench
- 57%
- AA Index
- N/A
- Context
- 205k
- Price in / out
- $0.3 / N/A per M
Leads SWE-Bench Verified (87.6) and Pro (64.3); new tokenizer (~35% more tokens). Claude Opus 4.7 supports vision input, tool calling, extended reasoning and computer use, with a 1M-token context window.
- SWE-Bench Verified
- 87.6%
- SWE-Bench Pro
- 64.3%
- Terminal-Bench
- 69.4%
- AA Index
- N/A
- Context
- 1M
- Price in / out
- $5 / $25 per M
754B · #1 SWE-Bench Pro among open weights. GLM-5.1 supports vision input, tool calling and extended reasoning, with a 128k-token context window.
- SWE-Bench Verified
- 79%
- SWE-Bench Pro
- 58.4%
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 128k
- Price in / out
- $0.11 / N/A per M
4-agent backbone · IFBench #1 (83); 8% of Opus cost. Grok 4.20 Beta 2 supports vision input and tool calling, with a 256k-token context window.
- SWE-Bench Verified
- 75%
- SWE-Bench Pro
- N/A
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 256k
- Price in / out
- $2 / N/A per M
Full SDLC agent: debug, terminal, PRDs, tests. GPT-5.3 Codex supports vision input, tool calling, extended reasoning and computer use, with a 400k-token context window.
- SWE-Bench Verified
- 80%
- SWE-Bench Pro
- 56.8%
- Terminal-Bench
- 77.3%
- AA Index
- N/A
- Context
- 400k
- Price in / out
- $10 / N/A per M
Adjustable Deep Think; agentic browsing (BrowseComp 85.9). Gemini 3.1 Pro supports vision input, audio, tool calling, extended reasoning and computer use, with a 2M-token context window.
- SWE-Bench Verified
- 78%
- SWE-Bench Pro
- N/A
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 2M
- Price in / out
- $7 / N/A per M
125B · code-specialized variant. Devstral-2-123B supports tool calling and computer use, with a 128k-token context window.
- SWE-Bench Verified
- N/A
- SWE-Bench Pro
- N/A
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 128k
- Price in / out
- $0.6 / N/A per M
Large-codebase reasoning, multi-file refactors. Claude Opus 4.6 supports vision input, tool calling, extended reasoning and computer use, with a 1M-token context window.
- SWE-Bench Verified
- 80.8%
- SWE-Bench Pro
- 57.3%
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 1M
- Price in / out
- $15 / N/A per M
675B MoE / 41B active · non-reasoning (AIME ~40, GPQA ~44) · Apache 2.0. Mistral Large 3 supports vision input and tool calling, with a 256k-token context window.
- SWE-Bench Verified
- N/A
- SWE-Bench Pro
- N/A
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 256k
- Price in / out
- $2 / N/A per M
Default coding driver. 98% of Opus at ⅕ cost. Claude Sonnet 4.6 supports vision input, tool calling, extended reasoning and computer use, with a 1M-token context window.
- SWE-Bench Verified
- 79.6%
- SWE-Bench Pro
- N/A
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 1M
- Price in / out
- $3 / $15 per M
4 to 5 times faster than Sonnet 4.5; cheap multi-agent driver. Claude Haiku 4.5 supports vision input, tool calling, extended reasoning and computer use, with a 200k-token context window.
- SWE-Bench Verified
- 73.3%
- SWE-Bench Pro
- N/A
- Terminal-Bench
- N/A
- AA Index
- N/A
- Context
- 200k
- Price in / out
- $1 / $5 per M
Insufficient information to confirm frontier-class capabilities; model card provides no substantive details beyond routing to a model family, making full evaluation impossible.
- AA Index
- Awaiting benchmarks
- Context
- 1.1M
- Price in / out
- $10 / $50 per M
- Found via
- OpenRouter
Insufficient information to assess the model's true capabilities, architecture, or scale; the model card only states it redirects to the latest GPT Sol family model with no further details.
- AA Index
- Awaiting benchmarks
- Context
- 1.1M
- Price in / out
- $2 / $10 per M
- Found via
- OpenRouter
Insufficient information to assess the model's true capabilities, architecture, or scale; the model card provides only a routing description with no technical details.
- AA Index
- Awaiting benchmarks
- Context
- 1.1M
- Price in / out
- $2 / $12 per M
- Found via
- OpenRouter
Insufficient information to determine if this is a frontier-class model; the model card provides almost no technical details, params are unknown, and 'GPT Luna' is not a recognized public OpenAI release as of the knowledge cutoff.
- AA Index
- Awaiting benchmarks
- Context
- 1.1M
- Price in / out
- $0.19999999999999998 / $1.2 per M
- Found via
- OpenRouter
Nex-N2.5-Pro is a ~397B MoE agentic model with multimodal (image+text) capabilities, computer use, and web browsing, representing a meaningful frontier-class post-training effort on a large-scale foundation.
- AA Index
- Awaiting benchmarks
- Context
- N/A
- Price in / out
- N/A / N/A per M
- Found via
- Hugging Face
Novel Causal Encoder-Decoder (CED) architecture from a major lab is significant, but small active parameter count (8B) limits scale score despite the architectural innovation.
- AA Index
- Awaiting benchmarks
- Context
- 1M
- Price in / out
- $0.15 / $0.6 per M
- Found via
- OpenRouter
GPT-6 Astra Pro appears to be a frontier-class model from OpenAI with a very large context window and high pricing suggesting significant scale, but it is merely a reasoning-mode variant of GPT-6 Astra rather than a distinct model, and no verifiable metadata (params, architecture, capabilities) is available to confirm claims.
- AA Index
- Awaiting benchmarks
- Context
- 1.1M
- Price in / out
- $10 / $50 per M
- Found via
- OpenRouter
GLM-5.3 is a 753B-parameter open-weights MoE model achieving frontier-class coding and cyber capability benchmarks, representing a significant post-training advancement over GLM-5.2 with SOTA results on multiple public benchmarks.
- AA Index
- Awaiting benchmarks
- Context
- N/A
- Price in / out
- N/A / N/A per M
- Found via
- Hugging Face
2.4T parameter MoE model with multimodal capabilities (text/image/video) and 1M context window represents frontier-class scale and capability.
- AA Index
- Awaiting benchmarks
- Context
- 1M
- Price in / out
- $2 / $6 per M
- Found via
- OpenRouter
Tencent's 780B MoE flagship model represents frontier-class scale and capability with a new-generation architecture from a major lab.
- AA Index
- Awaiting benchmarks
- Context
- N/A
- Price in / out
- N/A / N/A per M
- Found via
- Hugging Face
Qwen3.8 Flash appears to be a capable multimodal reasoning model with a 1M context window and broad capabilities, but parameter count is unknown and it may be a smaller/flash variant, keeping composite below include threshold.
- AA Index
- Awaiting benchmarks
- Context
- 1M
- Price in / out
- $0.15 / $0.47 per M
- Found via
- OpenRouter
GLM-5.3-Flash features a novel hybrid sparse/linear attention architecture with a very large 1.3M token context window suited for agent tasks, but parameter count is unknown and it's a flash/efficient variant rather than a frontier-scale model.
- AA Index
- Awaiting benchmarks
- Context
- 1.3M
- Price in / out
- $0.15 / $0.5 per M
- Found via
- OpenRouter
DeepSeek-V4-Flash-0731 is a frontier-class 304B-parameter MoE model from a major lab with a dedicated technical report, high download counts, and represents a new official release superseding a preview version.
- AA Index
- Awaiting benchmarks
- Context
- N/A
- Price in / out
- N/A / N/A per M
- Found via
- Hugging Face
What the columns mean
- SWE-Bench Verified
- Share of 500 real GitHub issues the model fixed end to end, tests passing. The best single signal for coding agents.
- SWE-Bench Pro
- A harder, contamination-resistant set of long-horizon coding tasks. Scores run lower than Verified for every model.
- Terminal-Bench
- Tasks completed in a real shell: builds, data wrangling, debugging. Measures agentic work, not chat.
- Reasoning
- Our composite of GPQA Diamond, Humanity's Last Exam, AIME, and ARC-AGI-2, normalised to 0 to 100. Use it to rank, not to quote.
- Context
- How much text fits in one request. 1M tokens is roughly 750,000 words. Usable context is often less than the claimed figure.
- $/M in
- US dollars per million input tokens at list price. Output tokens cost more, usually 3 to 5 times. Open models show the cheapest hosted rate.
- Capabilities
- V vision input, A audio, T tool calling, R extended reasoning mode, C computer or browser use. A missing letter means not stated, not no.
- AA Index
- The Artificial Analysis Intelligence Index: one independently measured number across reasoning, coding, and agentic evals. The only score on this page that every model, new or old, gets on the same day.
- Live rows
- Found nightly on Hugging Face and OpenRouter and admitted by a Claude judge. They carry no benchmark scores until the vendor publishes them.
Video generation models
| Model | Vendor | License | Released | I2V Rank | Max len (s) | Resolution | Audio | $/gen | Best for |
|---|---|---|---|---|---|---|---|---|---|
| Kling 3.0 | Kuaishou | closed | 2026-02 | 1 | 15 (multi-shot) | 1080p | synced dialogue + SFX | 1.00 | Cinematic shots; #1 general-purpose |
| Veo 3.1 | closed | 2026-01 | 2 | 8 | 1080p | yes | 0.75 | High-fidelity realism, prompt adherence | |
| Sora 2 | OpenAI | closed | 2025-10 | 3 | 12 | 1080p | yes | 0.90 | Imaginative T2V; ChatGPT-integrated |
| Seedance 2.0 | ByteDance | closed | 2026-04 | 4 | 15 | 1080p | yes | 0.50 | Product ads, e-comm, character consistency |
| LTX-2 | Lightricks | open | 2025-Q4 | 5 | 10 | 4K@50fps | native sync | 0.20 | Open-weights leader; 4K + audio |
| Wan 2.2 | Alibaba | open | 2025-Q3 | 6 | 6 | 720p | N/A | 0.05 | Runs on a 4070; novel MoE denoiser |
Image generation models
| Model | Vendor | License | Released | Max res | Quality* | Speed (s) | $/img | Best for |
|---|---|---|---|---|---|---|---|---|
| Midjourney v8 | Midjourney | closed | 2026-03 | 2K | 95 | 10 | 0.04 | Aesthetic leader; rewritten engine, five times faster than v7 |
| FLUX.2 [pro] | Black Forest Labs | closed (API) | 2026-Q1 | 2K | 93 | 4.5 | 0.04 | Photoreal commercial: best skin, lighting, materials |
| FLUX.2 [dev] | Black Forest Labs | open weights | 2026-Q1 | 2K | 89 | 6 | 0 | Top open-weights photoreal model |
| GPT Image 2 | OpenAI | closed | 2026-Q1 | 1K+ | 92 | 6 | 0.04 | Best prompt adherence: complex composed scenes |
| Imagen 4 Ultra | closed | 2025-Q4 | 2K | 94 | 5 | 0.04 | Photoreal flagship; strong text rendering | |
| Imagen 4 Fast | closed | 2026-03 | 1K | 86 | 2 | 0.02 | Cheapest fast quality at $0.02/img | |
| Nano Banana 2 | Google (Gemini 3.1 Flash Image) | closed | 2026-03 | 1K | 82 | 1.5 | 0.02 | Fastest end to end, 1 to 3 seconds; in-chat editing |
| Recraft V4 | Recraft | closed | 2025-Q4 | 2K | 88 | 5 | 0.04 | Brand / design assets, vector-style outputs |
| Ideogram 3 | Ideogram | closed | 2025-Q4 | 1K | 84 | 5 | 0.03 | Typography & in-image text rendering |
| Seedream 4.5 | ByteDance | closed | 2026-Q1 | 2K | 90 | 5 | 0.03 | Asian aesthetics, character consistency |
| Stable Diffusion 4 | Stability AI | open | 2026-Q1 | 2K | 80 | 6 | 0 | Open ecosystem; ControlNet and LoRA backbone |
| Hunyuan Image 3 | Tencent | open | 2026-Q1 | 2K | 83 | 5 | 0 | Open Tencent flagship; bilingual prompts |