Browse benchmarks
AllAgentsCodingConversationalGeneralKnowledgeKnowledge WorkLegalLong ContextMathMultimodalNLPReasoningSafetySearchSecurityTool UseTranslationVision
9 benchmarks in Knowledge
ARC
https://allenai.org/data/arcKnowledge1mo ago
C-Eval
https://cevalbenchmark.com/Knowledge1mo ago
FACTS Grounding
https://deepmind.google/discover/blog/facts-grounding-a-new-benchmark-for-evaluating-the-factuality-of-large-language-models/Knowledge1mo ago
MMLU
https://github.com/hendrycks/testKnowledge1mo ago
MMLU-Pro
https://huggingface.co/datasets/TIGER-Lab/MMLU-ProKnowledge1mo ago
MMMLU
Knowledge1mo ago
PubMedQA
https://pubmedqa.github.io/Knowledge1mo ago
SimpleQA
https://openai.com/index/introducing-simpleqa/Knowledge1mo ago
TriviaQA
https://nlp.cs.washington.edu/triviaqa/Knowledge1mo ago