Browse benchmarks
AllAgentsCodingConversationalGeneralKnowledgeKnowledge WorkLegalLong ContextMathMultimodalNLPReasoningSafetySearchSecurityTool UseTranslationVision
142 benchmarks
AA Coding Agent Index
Coding44m ago
AI2D
https://allenai.org/data/diagramsMultimodal1h ago
AIME
Math1h ago
ALERT
https://github.com/Babelscape/ALERTSafety1h ago
ANLI
https://github.com/facebookresearch/anliNLP1h ago
API-Bank
https://github.com/AlibabaResearch/DAMO-ConvAI/tree/main/api-bankAgents1h ago
APPS
https://github.com/hendrycks/appsCoding44m ago
ARC
https://allenai.org/data/arcKnowledge1h ago
ARC-AGI
https://arcprize.org/Reasoning1h ago
ARC-AGI 2
https://arcprize.org/Reasoning44m ago
ARC-AGI 3
https://arcprize.org/Reasoning44m ago
AgentBench
https://github.com/THUDM/AgentBenchAgents1h ago
Aider Polyglot
https://aider.chat/docs/leaderboards/Coding1h ago
AlpacaEval
https://github.com/tatsu-lab/alpaca_evalConversational1h ago
Arena-Hard
https://github.com/lm-sys/arena-hard-autoConversational1h ago
AutomationBench
Agents44m ago
BBH
https://github.com/suzgunmirac/BIG-Bench-HardReasoning1h ago
BBQ
https://github.com/nyu-mll/BBQSafety1h ago
BFCL
https://gorilla.cs.berkeley.edu/leaderboard.htmlCoding1h ago
BIG-Bench
https://github.com/google/BIG-benchReasoning1h ago
BigCodeBench
https://bigcode-bench.github.io/Coding1h ago
BigLaw Bench
Legal44m ago
BoolQ
https://github.com/google-research-datasets/boolean-questionsNLP1h ago
BrowseComp
Search44m ago
BrowseComp-Plus
Search44m ago
C-Eval
https://cevalbenchmark.com/Knowledge1h ago
COPA
https://people.ict.usc.edu/~gordon/copa.htmlReasoning1h ago
COSMOS QA
https://wilburone.github.io/cosmos/Reasoning1h ago
CRUXEval
https://github.com/facebookresearch/cruxevalCoding1h ago
ChartQA
https://github.com/vis-nlp/ChartQAMultimodal1h ago
Chatbot Arena
https://lmarena.ai/Conversational1h ago
ClassEval
https://github.com/FudanSELab/ClassEvalCoding1h ago
CoQA
https://stanfordnlp.github.io/coqa/NLP1h ago
CodeContests
https://github.com/google-deepmind/code_contestsCoding1h ago
CommonSenseQA
https://www.tau-nlp.sites.tau.ac.il/commonsenseqaReasoning1h ago
CursorBench
Coding44m ago
CursorBench 3.2
Coding44m ago
CyScenarioBench
Security44m ago
CyberGym
Security44m ago
DROP
https://allenai.org/data/dropNLP1h ago
DS-1000
https://ds1000-code-gen.github.io/Coding1h ago
DecodingTrust
https://decodingtrust.github.io/Safety1h ago
DeepSearchQA
Search44m ago
DocVQA
https://www.docvqa.org/Multimodal1h ago
EvalPlus
https://github.com/evalplus/evalplusCoding1h ago
FACTS Grounding
https://deepmind.google/discover/blog/facts-grounding-a-new-benchmark-for-evaluating-the-factuality-of-large-language-models/Knowledge1h ago
FLORES
https://github.com/facebookresearch/floresTranslation1h ago
FrontierBench
General44m ago
FrontierCode
Coding44m ago
FrontierCode 1.1
Coding44m ago
1 / 3next →