benchmarkdns
BrowseRegister

Browse benchmarks

AllAgentsCodingConversationalGeneralKnowledgeKnowledge WorkLegalLong ContextMathMultimodalNLPReasoningSafetySearchSecurityTool UseTranslationVision

7 benchmarks in Conversational

AlpacaEval
Conversational1h ago
https://github.com/tatsu-lab/alpaca_eval
Arena-Hard
Conversational1h ago
https://github.com/lm-sys/arena-hard-auto
Chatbot Arena
Conversational1h ago
https://lmarena.ai/
IFEval
Conversational1h ago
https://github.com/google-research/google-research/tree/master/instruction_following_eval
LiveBench
Conversational1h ago
https://livebench.ai/
MT-Bench
Conversational1h ago
https://github.com/lm-sys/FastChat/tree/main/fastchat/llm_judge
WildBench
Conversational1h ago
https://github.com/allenai/WildBench
benchmarkdns — a collision-avoidance registry for LLM benchmark names
Built by New Measure team