TASTE 基准测试模型判断 AI 安全研究提案的能力
TASTE tests whether models can judge AI-safety research proposals
Anthropic Fellows 发布 TASTE,以 92 组两两比较测试模型能否判断 AI 安全研究提案。团队估算人类研究者的一致率为 77%,表现最佳的 Fable 5 为 60%。这是规模较小的早期基准,置信区间较宽;结果表明模型仍落后于人类判断,而不是证明某个系统具备可靠的研究评审能力。
Anthropic Fellows released TASTE, a benchmark of 92 pairwise comparisons asking models to judge AI-safety research proposals. The team estimates human researcher agreement at 77%, while the best model, Fable 5, reached 60%. This is a small early benchmark with wide confidence intervals; it suggests a gap from human judgment rather than establishing reliable automated peer review.