Browse benchmarks
AllAgentsCodingConversationalGeneralKnowledgeKnowledge WorkLegalLong ContextMathMultimodalNLPReasoningSafetySearchSecurityTool UseTranslationVision
7 benchmarks in Conversational
AlpacaEval
https://github.com/tatsu-lab/alpaca_evalConversational1h ago
Arena-Hard
https://github.com/lm-sys/arena-hard-autoConversational1h ago
Chatbot Arena
https://lmarena.ai/Conversational1h ago
IFEval
https://github.com/google-research/google-research/tree/master/instruction_following_evalConversational1h ago
LiveBench
https://livebench.ai/Conversational1h ago
MT-Bench
https://github.com/lm-sys/FastChat/tree/main/fastchat/llm_judgeConversational1h ago
WildBench
https://github.com/allenai/WildBenchConversational1h ago