ARTICLE · INTELLIGENCE

战地情报 · 详情页

来自尧图项目组的一线实战观察与深度解析

吐血整理大模型评测信息(含三方测评发布网站),动动小手收藏备用~

吐血整理大模型评测信息(含三方测评发布网站),动动小手收藏备用~ AA Intelligence Index Artificial AnalysisLeaderboard:https://artificialanalysis.ai/leaderboards/modelshttps://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index官网https://artificialanalysis.ai简介Artificial Analysis 基于独立复测的多项基准含 GPQA、IFBench、Terminal-Bench、SciCode 等加权合成的智能指数当前版本 v4.1.1 已收录 181 个模型Claude Opus 5、Claude Fable 5、Grok 4.6 等居前因口径统一、更新及时是业界引用最多的模型横向能力排名之一ALE Agents’ Last ExamLeaderboard https://snorkel.ai/leaderboard/agents-last-exam/官网https://agents-last-exam.org论文https://arxiv.org/html/2606.05405v1简介由 Berkeley RDI 牵头、300 行业专家共建的超大规模真实工作 Agent 基准目标 5,000 个任务、覆盖 55 个子行业衡量长程、高经济价值、可验证结果的职业任务完成能力是当前评估AI 能否胜任真实数字劳动的代表性基准HLEHumanity’s Last ExamLeaderboardhttps://lastexam.ai/官网https://lastexam.ai/论文https://arxiv.org/pdf/2501.14249简介由 Center for AI Safety 与 Scale AI 发起、全球学科专家出题的人类终极考试覆盖 100 学科的超难前沿题用于逼近人类知识边界、区分顶级模型在饱和基准之外的极限能力是公认的学术天花板测试Terminal-BenchLeaderboard 2.0官方: https://www.tbench.ai/leaderboard/terminal-bench/2.0三方https://www.datalearner.com/benchmarks/terminalbench-2Leaderboard 2.1官方https://www.tbench.ai/leaderboard/terminal-bench/2.1三方https://www.datalearner.com/benchmarks/terminal-bench-2-1Leaderboard 3.0官方https://www.frontierbench.ai/三方https://www.datalearner.com/benchmarks/terminal-bench-v3官网https://www.tbench.ai简介由 Sierra 联合多家机构推出的终端命令行环境长程任务基准2.1 版修正了 2.0 的 28 个任务并引入连续验证机制用于衡量 Agent 在真实终端工作流中的复杂问题解决能力是 CLI Agent 领域的代表性基准AutomationBenchLeaderboardhttps://zapier.com/benchmarks官网https://github.com/zapier/AutomationBench论文https://arxiv.org/html/2604.18934v1简介Zapier 推出的跨应用工作流编排基准任务覆盖销售、营销、运营、客服、财务、HR 六大业务域Agent 需自主发现 API 端点、遵循分层业务规则并正确写入各系统用于衡量真实业务自动化执行这一企业级能力官方榜单采用未公开的私有任务集以防过拟合SWE-Bench VerifiedLeaderboard:官方 https://www.swebench.com/三方https://llm-stats.com/benchmarks/swe-bench-verified官网https://openai.com/index/introducing-swe-bench-verified/论文https://arxiv.org/pdf/2310.06770简介从 12 个热门 Python 仓库的真实 issue 与 PR 中抽取、经人工校验的 500 题子集是目前编码 Agent 领域被引用最多、事实上的编码能力标尺可作为横向比较代码 Agent 的首选参考SWE-Bench ProLeaderboard: https://www.datalearner.com/benchmarks/swe-bench-pro官网https://github.com/scaleapi/SWE-bench_Pro-os简介Scale AI 为解决数据污染、任务单一、问题过简、测试不可复现四大痛点而设计的进阶基准覆盖 41 个专业仓库、1,865 个真实任务前沿模型在标准化脚手架下仅能解出约 59%是衡量去污染真实工程能力的更严苛参考DeepSWELeaderboardhttps://deepswe.datacurve.ai/官网https://deepswe.datacurve.ai论文https://arxiv.org/abs/2607.07946简介由 Datacurve非 DeepLearning.ai推出的 113 个原创长程软件工程任务基准任务从零编写、从不回馈上游仓库从源头规避预训练污染并以手写验证器替代随修复附带的测试是衡量无污染真实编码能力的重要参考NL2Repo-BenchLeaderboardhttps://www.datalearner.com/benchmarks/nl2repo-bench论文https://arxiv.org/abs/2512.12730简介由 M-A-P、ByteDance Seed 等机构联合提出的仓库级代码生成基准——仅凭一份自然语言需求文档和空工作区Agent 需自主完成架构设计、依赖管理与多模块实现并产出可安装的 Python 库用于评估当前最难的长程仓库级生成能力目前最强 Agent 平均测试通过率仍不足 40%SciCodeLeaderboardhttps://pricepertoken.com/leaderboards/benchmark/scicode官网https://scicode-bench.github.io论文https://arxiv.org/abs/2407.13168简介由 16 个自然科学领域的科学家联合策划的科研代码基准含 80 个主问题、338 个子问题考察知识回忆推理代码综合的复合能力是学术界公认的科研编码难度标尺NeurIPS 2024 DBCyberGymLeaderboardhttps://benchlm.ai/benchmarks/cybergym官网https://github.com/sunblaze-ucb/cybergym论文https://arxiv.org/html/2606.04460v2简介UC BerkeleyDawn Song 组推出的规模化网络安全基准覆盖 188 个广泛使用的开源项目如 OpenSSL、FFmpeg中的 1,507 个真实 CVE采用执行式客观评估是衡量 AI 白盒漏洞分析与安全工程能力的权威参考PaperBenchLeaderboardhttps://github.com/openai/frontier-evals/tree/main/project/paperbench官网https://github.com/openai/frontier-evals/tree/main/project/paperbench论文https://arxiv.org/abs/2504.01848简介OpenAI 推出的科研复现基准要求 Agent 从零复现 20 篇 ICML 2024 Spotlight/Oral 论文共 8,316 个可评分子项与原论文作者共同制定 rubric最佳 Agent 平均复现得分仅约 21%是衡量端到端科研工程能力目前最难公开基准之一IFBenchLeaderboard官方https://benchlm.ai/benchmarks/ifbench三方https://www.datalearner.com/benchmarks/if-benchhttps://artificialanalysis.ai/evaluations/ifbenchhttps://benchlm.ai/benchmarks/ifbench论文https://arxiv.org/pdf/2507.02833简介Allen AIAi2推出的精确指令遵循泛化基准含 58 个全新可验证的 OOD 输出约束专测模型对没见过的新约束的遵循能力已被 Artificial Analysis 纳入综合智能指数是评估听话程度泛化性的主流参考LiveCodeBenchLeaderboardhttps://livecodebench.github.io/leaderboard.html官网https://livecodebench.github.iohttps://github.com/LiveCodeBench/LiveCodeBench简介持续从 LeetCode、AtCoder、Codeforces 收集新题的无污染竞赛编程基准并额外考察自修复、代码执行、测试输出预测等能力通过按题目发布日期切分实现防污染评测是评估实时编码能力的首选参考OSWorld VerifiedLeaderboardhttps://osworld-v1.xlang.ai官网https://github.com/xlang-ai/OSWorldVerified: https://xlang.ai/blog/osworld-verified论文https://proceedings.neurips.cc/paper_files/paper/2024/hash/5d413e48f84dc61244b6be550f1cd8f5-Abstract-Datasets_and_Benchmarks_Track.html简介HKU XLANG Lab 推出的首个真实计算机环境Windows/Ubuntu/macOS多模态 GUI Agent 基准Verified 版对 369 个任务做了人工校验是电脑操作 AgentComputer Use领域最权威的参考进阶可关注长程版 OSWorld 2.0GPQA DiamondLeaderboardEpoch AIhttps://epoch.ai/benchmarks/gpqa-diamondArtificial Analysishttps://artificialanalysis.ai/evaluations/gpqa-diamondVals AIhttps://www.vals.ai/benchmarks/gpqaLLM Statshttps://llm-stats.com/benchmarks/gpqa llm-stats.com论文https://arxiv.org/pdf/2311.12022简介由学科专家多为博士撰写的 198 道Google-proof研究生级理科多选题博士专家本人仅约 69.7% 正确率目前已有 24 个模型超过 90%属明显饱和的知识推理参考仍可作为顶级模型理科推理的快速筛选指标参考位置三方整合榜单AAArtificial Analysishttps://artificialanalysis.ai/leaderboards/modelsVals.AIhttps://www.vals.ai/benchmarksLLM STATS benchmarkshttps://llm-stats.com/benchmarksEpoch.aihttps://epoch.ai/benchmarksSnorkel benckmark leaderboardshttps://snorkel.ai/leaderboard/Dataleaner Benchmarkshttps://www.datalearner.com/benchmarksBenchLM benchmarkshttps://benchlm.ai/benchmarks
RELATED READING

延伸阅读

更多一线实战笔记与深度复盘,助您持续精进