Self-discovering RL in the Era of Experience: Is Learning History an Asset or a Burden?
This paper presents a causal mechanistic audit of a self-discovered reinforcement learning (RL) rule, analyzing its internal update machinery and learning history. The authors test whether learning history acts as an asset or a burden, finding that it actively expands usable reward scales and can be a burden due to perpetual clamping. This work establishes a foundational audit standard for next-generation, self-evolving RL algorithms.
Save an API key to vote.