Browse benchmarks
AllAgentsCodingConversationalGeneralKnowledgeKnowledge WorkLegalLong ContextMathMultimodalNLPReasoningSafetySearchSecurityTool UseTranslationVision
10 benchmarks in Multimodal
AI2D
https://allenai.org/data/diagramsMultimodal1h ago
ChartQA
https://github.com/vis-nlp/ChartQAMultimodal1h ago
DocVQA
https://www.docvqa.org/Multimodal1h ago
MM-Vet
https://github.com/yuweihao/MM-VetMultimodal1h ago
MMBench
https://github.com/open-compass/MMBenchMultimodal1h ago
MME
https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models/tree/EvaluationMultimodal1h ago
MMMU
https://mmmu-benchmark.github.io/Multimodal1h ago
RealWorldQA
https://huggingface.co/datasets/xai-org/RealWorldQAMultimodal1h ago
SEED-Bench
https://github.com/AILab-CVC/SEED-BenchMultimodal1h ago
VQAv2
https://visualqa.org/Multimodal1h ago