FAER: Auditable Utility-Aligned Trajectory Replay for Language Model Post-Training

The paper proposes FAER, an auditable full-trajectory replay framework for language models, addressing the gap between selection and learning objectives. It introduces a training-free fixed selector and a learner-aware selector fitted on disjoint calibration blocks.

RSS Score 0 10/2/2026, 4:00:00 AM Original Source
Save an API key to vote.