AI Model Leaderboard

10 LLMs ranked by benchmark performance & API pricing. Updated daily.

Full Benchmark Breakdown

All scores from public leaderboards. Higher is better.

ModelMMLUHumanEvalSWE-benchGSM8KGPQAIFEvalELO
Gemini90%88.4%36.1%95.8%62.2%84.1%1301
ChatGPT88.7%90.2%33.2%95.8%53.6%85.6%1287
Claude89.3%93.7%49%96.4%59.4%89.3%1271
Grok86%85%30%94%48%1268
DeepSeek88.5%82.6%38.8%97.3%51.1%78.3%1257
Llama 486%81%28%94%45%1228
Tongyi Qianwen85%80%93%44%1220
Mistral Large81%81%24%90%38%1215
GLM-584%78%92%42%1205
ϕPhi-479%76%88%35%1185

Compare models side-by-side

Use our interactive comparison deck with weighted scoring.

Compare AI Tools