Browse benchmarks
AllAgentsCodingConversationalGeneralKnowledgeKnowledge WorkLegalLong ContextMathMultimodalNLPReasoningSafetySearchSecurityTool UseTranslationVision
142 benchmarks
AA Coding Agent Index
Coding18m ago
AI2D
https://allenai.org/data/diagramsMultimodal57m ago
AIME
Math57m ago
ALERT
https://github.com/Babelscape/ALERTSafety57m ago
ANLI
https://github.com/facebookresearch/anliNLP57m ago
API-Bank
https://github.com/AlibabaResearch/DAMO-ConvAI/tree/main/api-bankAgents57m ago
APPS
https://github.com/hendrycks/appsCoding18m ago
ARC
https://allenai.org/data/arcKnowledge57m ago
ARC-AGI
https://arcprize.org/Reasoning57m ago
ARC-AGI 2
https://arcprize.org/Reasoning18m ago
ARC-AGI 3
https://arcprize.org/Reasoning18m ago
AgentBench
https://github.com/THUDM/AgentBenchAgents57m ago
Aider Polyglot
https://aider.chat/docs/leaderboards/Coding57m ago
AlpacaEval
https://github.com/tatsu-lab/alpaca_evalConversational57m ago
Arena-Hard
https://github.com/lm-sys/arena-hard-autoConversational57m ago
AutomationBench
Agents18m ago
BBH
https://github.com/suzgunmirac/BIG-Bench-HardReasoning57m ago
BBQ
https://github.com/nyu-mll/BBQSafety57m ago
BFCL
https://gorilla.cs.berkeley.edu/leaderboard.htmlCoding57m ago
BIG-Bench
https://github.com/google/BIG-benchReasoning57m ago
BigCodeBench
https://bigcode-bench.github.io/Coding57m ago
BigLaw Bench
Legal18m ago
BoolQ
https://github.com/google-research-datasets/boolean-questionsNLP57m ago
BrowseComp
Search18m ago
BrowseComp-Plus
Search18m ago
C-Eval
https://cevalbenchmark.com/Knowledge57m ago
COPA
https://people.ict.usc.edu/~gordon/copa.htmlReasoning57m ago
COSMOS QA
https://wilburone.github.io/cosmos/Reasoning57m ago
CRUXEval
https://github.com/facebookresearch/cruxevalCoding57m ago
ChartQA
https://github.com/vis-nlp/ChartQAMultimodal57m ago
Chatbot Arena
https://lmarena.ai/Conversational57m ago
ClassEval
https://github.com/FudanSELab/ClassEvalCoding57m ago
CoQA
https://stanfordnlp.github.io/coqa/NLP57m ago
CodeContests
https://github.com/google-deepmind/code_contestsCoding57m ago
CommonSenseQA
https://www.tau-nlp.sites.tau.ac.il/commonsenseqaReasoning57m ago
CursorBench
Coding18m ago
CursorBench 3.2
Coding18m ago
CyScenarioBench
Security18m ago
CyberGym
Security18m ago
DROP
https://allenai.org/data/dropNLP57m ago
DS-1000
https://ds1000-code-gen.github.io/Coding57m ago
DecodingTrust
https://decodingtrust.github.io/Safety57m ago
DeepSearchQA
Search18m ago
DocVQA
https://www.docvqa.org/Multimodal57m ago
EvalPlus
https://github.com/evalplus/evalplusCoding57m ago
FACTS Grounding
https://deepmind.google/discover/blog/facts-grounding-a-new-benchmark-for-evaluating-the-factuality-of-large-language-models/Knowledge57m ago
FLORES
https://github.com/facebookresearch/floresTranslation57m ago
FrontierBench
General18m ago
FrontierCode
Coding18m ago
FrontierCode 1.1
Coding18m ago
1 / 3next →