Browse benchmarks
AllAgentsCodingConversationalGeneralKnowledgeKnowledge WorkLegalLong ContextMathMultimodalNLPReasoningSafetySearchSecurityTool UseTranslationVision
142 benchmarks
FrontierMath
https://epochai.org/frontiermathMath1h ago
GAIA
https://huggingface.co/gaia-benchmarkAgents1h ago
GDPval-AA
Knowledge Work32m ago
GDPval-AA v2
Knowledge Work32m ago
GLUE
https://gluebenchmark.com/NLP1h ago
GPQA
https://github.com/idavidrein/gpqaReasoning1h ago
GPQA-Diamond
https://github.com/idavidrein/gpqaReasoning1h ago
GSM8K
https://github.com/openai/grade-school-mathMath1h ago
HELM
https://crfm.stanford.edu/helm/Safety1h ago
HLEAutomationBench
Agents32m ago
HarmBench
https://github.com/centerforaisafety/HarmBenchSafety1h ago
HellaSwag
https://rowanzellers.com/hellaswag/Reasoning1h ago
HotpotQA
https://hotpotqa.github.io/NLP1h ago
HumanEval
https://github.com/openai/human-evalCoding1h ago
Humanity's Last Exam
https://lastexam.ai/Reasoning1h ago
IFEval
https://github.com/google-research/google-research/tree/master/instruction_following_evalConversational1h ago
InfiniteBench
https://github.com/OpenBMB/InfiniteBenchLong Context1h ago
LAMBADA
https://zenodo.org/records/2630551NLP1h ago
LiveBench
https://livebench.ai/Conversational1h ago
LiveCodeBench
https://livecodebench.github.io/Coding1h ago
LongBench
https://github.com/THUDM/LongBenchLong Context1h ago
MATH
https://github.com/hendrycks/mathMath1h ago
MBPP
https://github.com/google-research/google-research/tree/master/mbppCoding1h ago
MCP-Atlas
Tool Use32m ago
MGSM
https://github.com/google-research/url-nlp/tree/main/mgsmMath1h ago
MM-Vet
https://github.com/yuweihao/MM-VetMultimodal1h ago
MMBench
https://github.com/open-compass/MMBenchMultimodal1h ago
MME
https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models/tree/EvaluationMultimodal1h ago
MMLU
https://github.com/hendrycks/testKnowledge1h ago
MMLU-Pro
https://huggingface.co/datasets/TIGER-Lab/MMLU-ProKnowledge1h ago
MMMLU
Knowledge32m ago
MMMU
https://mmmu-benchmark.github.io/Multimodal1h ago
MRCR v2
Long Context32m ago
MT-Bench
https://github.com/lm-sys/FastChat/tree/main/fastchat/llm_judgeConversational1h ago
MathQA
https://math-qa.github.io/Math1h ago
MathVista
https://mathvista.github.io/Math1h ago
MuSR
https://github.com/Zayne-sprague/MuSRReasoning1h ago
MultiNLI
https://cims.nyu.edu/~sbowman/multinli/NLP1h ago
MultiPL-E
https://nuprl.github.io/MultiPL-E/Coding1h ago
Natural Questions
https://ai.google.com/research/NaturalQuestionsNLP1h ago
Needle in a Haystack
https://github.com/gkamradt/LLMTest_NeedleInAHaystackLong Context1h ago
NumGLUE
https://github.com/allenai/numglueMath1h ago
OSS-Fuzz
https://google.github.io/oss-fuzz/Security32m ago
OSWorld
https://os-world.github.io/Agents1h ago
OSWorld 2.0
Agents32m ago
OSWorld-Verified
Agents32m ago
OfficeQA Pro
Knowledge Work32m ago
OpenBookQA
https://allenai.org/data/open-book-qaReasoning1h ago
OpenRCA
Agents32m ago
PIQA
https://yonatanbisk.com/piqa/Reasoning1h ago