Browse benchmarks
AllAgentsCodingConversationalGeneralKnowledgeKnowledge WorkLegalLong ContextMathMultimodalNLPReasoningSafetySearchSecurityTool UseTranslationVision
9 benchmarks in Knowledge
ARC
https://allenai.org/data/arcKnowledge1h ago
C-Eval
https://cevalbenchmark.com/Knowledge1h ago
FACTS Grounding
https://deepmind.google/discover/blog/facts-grounding-a-new-benchmark-for-evaluating-the-factuality-of-large-language-models/Knowledge1h ago
MMLU
https://github.com/hendrycks/testKnowledge1h ago
MMLU-Pro
https://huggingface.co/datasets/TIGER-Lab/MMLU-ProKnowledge1h ago
MMMLU
Knowledge32m ago
PubMedQA
https://pubmedqa.github.io/Knowledge32m ago
SimpleQA
https://openai.com/index/introducing-simpleqa/Knowledge1h ago
TriviaQA
https://nlp.cs.washington.edu/triviaqa/Knowledge1h ago