Browse benchmarks
AllAgentsCodingConversationalGeneralKnowledgeKnowledge WorkLegalLong ContextMathMultimodalNLPReasoningSafetySearchSecurityTool UseTranslationVision
10 benchmarks in Multimodal
AI2D
https://allenai.org/data/diagramsMultimodal1mo ago
ChartQA
https://github.com/vis-nlp/ChartQAMultimodal1mo ago
DocVQA
https://www.docvqa.org/Multimodal1mo ago
MM-Vet
https://github.com/yuweihao/MM-VetMultimodal1mo ago
MMBench
https://github.com/open-compass/MMBenchMultimodal1mo ago
MME
https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models/tree/EvaluationMultimodal1mo ago
MMMU
https://mmmu-benchmark.github.io/Multimodal1mo ago
RealWorldQA
https://huggingface.co/datasets/xai-org/RealWorldQAMultimodal1mo ago
SEED-Bench
https://github.com/AILab-CVC/SEED-BenchMultimodal1mo ago
VQAv2
https://visualqa.org/Multimodal1mo ago