Repositories list
49 repositories
PowerBench
PublicCNFinBench
PublicAgentCompass
Public[EMNLP 2026] AgentCompass is an extensible open-source evaluation infrastructure for systematically assessing LLM/VLM agent capabilities.VLMEvalKit
PublicOpen-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarksopencompass
PublicOpenCompass is an LLM evaluation platform, supporting a wide range of models from OpenAI, Anthropic, Gemini, Qwen, GLM, DeepSeek, etc, across 100+ datasets cove…MCPServer-AC
Publicdocs
PublicGTA
Public[NeurIPS 2024 D&B] GTA: A Benchmark for General Tool Agents & [arXiv 2026] GTA-2GenEditEvalKit
PublicTextEdit
PublicMiroFlow
PublicRePro
Public[ICLR 2026] Rectifying LLM Thought From Lens of OptimizationSAGA
PublicATLAS
PublicOASIS
PublicInteractScience
PublicCognitiveKernel-Pro
PublicGAOKAO-Eval
Public.github
PublicMMBench-GUI
PublicOfficial repo of "MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents". It can be used to evaluate a GUI agent with a hierarchical mann…ReasonZoo
PublicCompassVerifier
PublicGPassK
Public[ACL 2025] Are Your LLMs Capable of Stable Reasoning?Creation-MMBench
PublicCompassJudger
PublicRaML
PublicBotChat
PublicAda-LEval
PublicThe official implementation of "Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks"MathBench
PublicMMBench
Public
ProTip! When viewing an organization's repositories, you can use the
props. filter to filter by custom property.