LLMs
1–25 of 420
Rank reflects Intelligence Index
| RankPosition based on the Intelligence Index. Tied scores share the same rank. | ModelModel name and publisher matched to the published benchmark result. | Composite score across reasoning, knowledge, mathematics, science, and coding evaluations. | Composite score for programming and software-engineering capability. | Composite score for tool use, planning, and autonomous task completion. | Maximum number of tokens the model can process in one request. | Median output generation throughput, measured in tokens per second. | Blended USD cost per 1 million tokens using a 3:1 input-to-output ratio. | Public release month and year for this model. | Whether downloadable model weights are open or access is closed. |
|---|---|---|---|---|---|---|---|---|---|
| 1 | 63.1 | 78.0 | 59.2 | 1M | 59 tokens/s | $10 | Jul 2026 | Closed source | |
| 2 | 62.1 | 76.5 | 56.6 | 1M | 71 tokens/s | $20 | Jun 2026 | Closed source | |
| 3 | 60.9 | 78.3 | 57.8 | 1M | 71 tokens/s | $11.25 | Jul 2026 | Closed source | |
| 4 | 60.9 | 76.8 | 58.7 | 500K | 69 tokens/s | $3 | Aug 2026 | Closed source | |
| 5 | 59.7 | 76.2 | 54.3 | 1M | 38 tokens/s | $6 | Jul 2026 | Open source | |
| 6 | 59.5 | 74.8 | 59.1 | 1M | 93 tokens/s | $2.15 | Aug 2026 | Closed source | |
| 7 | 58.1 | 71.8 | 58.4 | 1M | 45 tokens/s | $3 | Aug 2026 | Closed source | |
| 8 | 57.7 | 71.9 | 57.1 | 983.6K | 45 tokens/s | $3 | Aug 2026 | Open source | |
| 9 | 57.3 | 74.3 | 49.4 | 1M | 60 tokens/s | $10 | May 2026 | Closed source | |
| 10 | 56.8 | 72.2 | 49.3 | 1M | — | $2 | Aug 2026 | Closed source | |
| 11 | 56.6 | 76.7 | 50.2 | 1M | 116 tokens/s | $4.5 | Jul 2026 | Closed source | |
| 12 | 56.3 | 74.9 | 47.4 | 922K | 82 tokens/s | $11.25 | Apr 2026 | Closed source | |
| 13 | 56.0 | 76.1 | 45.1 | 1M | 394 tokens/s | $1.5 | Aug 2026 | Closed source | |
| 14 | 55.8 | 72.4 | 48.9 | 500K | 64 tokens/s | $3 | Jul 2026 | Closed source | |
| 15 | 55.3 | 71.5 | 49.7 | 1M | 81 tokens/s | $4 | Jun 2026 | Closed source | |
| 16 | 55.0 | 73.6 | 46.3 | 1M | 52 tokens/s | $10 | Apr 2026 | Closed source | |
| 17 | 53.2 | 71.3 | 39.7 | 1M | 270 tokens/s | $2 | Jul 2026 | Closed source | |
| 18 | 53.2 | 68.8 | 49.6 | 1M | 81 tokens/s | $1.98 | Aug 2026 | Open source | |
| 19 | 53.1 | 71.1 | 44.2 | 1.1M | 128 tokens/s | $5.63 | Mar 2026 | Closed source | |
| 20 | 52.6 | 68.8 | 45.7 | 1M | 95 tokens/s | $2.15 | Jun 2026 | Open source | |
| 21 | 52.3 | 71.4 | 46.9 | 1M | 157 tokens/s | $0.45 | Jul 2026 | Closed source | |
| 22 | 52.0 | 68.1 | 50.9 | 256K | 57 tokens/s | $1.13 | Aug 2026 | Open source | |
| 23 | 52.0 | 70.1 | 39.7 | 1M | 227 tokens/s | $3.38 | May 2026 | Closed source | |
| 24 | 51.8 | 69.1 | 48.4 | 1M | 137 tokens/s | $0.66 | Jul 2026 | Open source | |
| 25 | 51.6 | 69.2 | 40.5 | 1M | 228 tokens/s | $1.5 | Jul 2026 | Closed source | |
Rank uses Intelligence Index. Sort any index, context, speed, pricing, release date, or license; missing data always appears last. | |||||||||
Explore every published snapshot across coding, agentic search, reasoning, instruction following, long context, and safety evaluations.
Browse by category
24 of 24 benchmarks
A long-horizon repository benchmark built to test realistic software engineering work. Its scores need a task-quality and setup check before they support a coding-agent decision.
63 models
A multilingual software-engineering benchmark for real-world code issue resolution across multiple programming languages.
36 models
An independently evaluated Tau3 banking benchmark from Artificial Analysis.
16 models
How important are skills for agents?
29 models
A broad research-agent benchmark for open-ended information gathering, synthesis, and answer construction across wide search spaces.
14 models
A benchmark for web-browsing agents that must search, inspect sources, gather evidence, and return the correct answer to research-oriented questions.
40 models
MRCR v2 slice focused on very long contexts at 128K-256K lengths.
2 models
Human-validated software engineering issues from real-world Python repositories.
33 models
Factoid questions spanning politics, science, technology, art, sports, geography, and music.
78 models
Abstract reasoning and pattern generalization on grid-based tasks.
184 models
Exceptionally difficult research-level mathematics problems.
51 models
Real-world cybersecurity evaluation of AI agents reproducing vulnerabilities with working proof-of-concept tests.
51 models
Composite Artificial Analysis index across mathematics, science, coding, and reasoning evaluations.
603 models
Artificial Analysis composite index for programming and software-engineering capability.
228 models
Artificial Analysis composite index for tool use, planning, autonomy, and complex agentic workflows.
170 models
Artificial Analysis factual-knowledge and hallucination-resistance evaluation.
81 models
Graduate-level, expert-written science questions evaluated by Artificial Analysis.
584 models
Broad expert-level reasoning and knowledge evaluation published by Artificial Analysis.
575 models
Instruction-following benchmark results published by Artificial Analysis.
450 models
Scientific programming benchmark results published by Artificial Analysis.
575 models
Agentic terminal and software-engineering results published by Artificial Analysis.
208 models
Critical-points programming evaluation published by Artificial Analysis.
490 models
Long-context reasoning results published by Artificial Analysis.
508 models
Tool-agent interaction benchmark results published by Artificial Analysis.
440 models
Scores are published snapshots and are only comparable within the same benchmark, version, metric, and evaluation configuration.