10 LLMs ranked by benchmark performance & API pricing. Updated daily.
Scores: LMArena Chatbot Arena ELO (higher = better). Pricing: per 1M input tokens from OpenRouter API (auto-updated daily). How we score →
All scores from public leaderboards. Higher is better.
| Model | MMLU | HumanEval | SWE-bench | GSM8K | GPQA | IFEval | ELO |
|---|---|---|---|---|---|---|---|
| ✧Gemini | 90% | 88.4% | 36.1% | 95.8% | 62.2% | 84.1% | 1301 |
| ✦ChatGPT | 88.7% | 90.2% | 33.2% | 95.8% | 53.6% | 85.6% | 1287 |
| ✸Claude | 89.3% | 93.7% | 49% | 96.4% | 59.4% | 89.3% | 1271 |
| ✕Grok | 86% | 85% | 30% | 94% | 48% | — | 1268 |
| ◈DeepSeek | 88.5% | 82.6% | 38.8% | 97.3% | 51.1% | 78.3% | 1257 |
| ◑Llama 4 | 86% | 81% | 28% | 94% | 45% | — | 1228 |
| ◈Tongyi Qianwen | 85% | 80% | — | 93% | 44% | — | 1220 |
| ◆Mistral Large | 81% | 81% | 24% | 90% | 38% | — | 1215 |
| ▣GLM-5 | 84% | 78% | — | 92% | 42% | — | 1205 |
| ϕPhi-4 | 79% | 76% | — | 88% | 35% | — | 1185 |
Use our interactive comparison deck with weighted scoring.
Compare AI Tools