IBM 为 ALTK-Evolve 增加一致性分析,关注重复运行是否都能成功
IBM adds consistency analysis to ALTK-Evolve for repeatable agent success
IBM 团队介绍从已有执行轨迹中找出容易变化的决策点,再生成可复用指导的方法。在其 GPT-4.1 与 AppWorld 实验中,五次全部成功的任务比例从 53.0% 升至 69.0%。这是团队自测;Pass⁵ 要求每次都成功,不等于至少一次成功的 Pass@5,也不代表生产环境保证。
IBM describes identifying unstable decisions in recorded agent traces and turning them into reusable guidelines. In its GPT-4.1/AppWorld experiment, the share of tasks successful in all five runs rose from 53.0% to 69.0%. These are team-run results. Pass⁵ requires every run to succeed, unlike at-least-once Pass@5, and is not a production guarantee.