When Context Changes: Understanding Update Failures in LLMs
A research paper introduces Controlled In-Context Memory (CICM), a benchmark for tracking and using updated information in conversations and agent logs. The study finds that even frontier reasoning models can fail to recover the current state, and attention drift is identified as a mechanism for this failure. The paper proposes a solution to correct old-value errors without retraining models.
Save an API key to vote.