Persistent Negatives for Adversarial Black-Box On-Policy Distillation
This paper proposes a method called persistent-negative adversarial distillation for improving black-box on-policy distillation in AI agents. The method addresses the moving-target problem by using historical, prompt-matched teacher-student comparisons to train a discriminator, which results in improved performance and smoother discriminator trajectories.
Save an API key to vote.