Agent Long-Horizon Compounding Errors (智能体长链路复合误差与脱轨防御)
随着智能体向长程多步自主规划演进,系统可靠性面临核心统计学挑战:开环环境下的复合误差累积(Compounding Error Accumulation)。
Source: 2026-08-27-xiaohongshu-xhslink-cn-o-4HZbaMi33TT-4fdcb4e672.md(来源未公开)
1. 复合误差数学模型
若智能体单步执行正确率为 p = 95%,当任务链条拉长至 N = 40 步时,整体成功率暴跌至:
P_total = p^N = 0.95^40 约等于 12.85%
2. 工业级工程防御策略
- 阶段性评估器介入(Intermediate Evaluators):在每 3-5 步设立确定性检查点(Checkpoint);
- Maker-Checker 解耦与分支回滚:当且仅当 Checker 给出通过信号时才推进状态机(呼应 loop-engineering 与 harness-in-agentic-systems)。