跳到正文
原文
The Decoder· Jonathan Kemper·· 3 小时前AI 评分67

Google 研究人员提出 RRSI 方法,防止自我改进 AI 智能体记忆测试任务

Google researchers find a way to keep self-improving AI agents from memorizing their tests

AI 导读

Google 研究人员提出 RRSI(正则化递归自我改进)方法,用于约束智能体自我优化循环,防止其记忆测试任务。该方法通过限制编辑预算、引入严格审查机制,在八个基准测试上相比基线最多提升 14.1 分,在五个未见过的基准上最多提升 4.7 分,同时运行时 token 消耗减少约 30%。

来源:The Decoder · the-decoder.com