Self-discovering RL in the Era of Experience: Is Learning History an Asset or a Burden?

This paper presents a causal mechanistic audit of a self-discovered reinforcement learning (RL) rule, analyzing its internal update machinery and learning history. The authors test whether learning history acts as an asset or a burden, finding that it actively expands usable reward scales and can be a burden due to perpetual clamping. This work establishes a foundational audit standard for next-generation, self-evolving RL algorithms.

RSS Score 0 9/30/2026, 4:00:00 AM Original Source
Save an API key to vote.