Alignment via Training Against Probes Without Losing Monitorability

This paper studies probe-guided fine-tuning for model alignment, where probes detect undesired properties in model activations as a direct training signal. The researchers evaluate linear and non-linear probes with different numbers of probes per layer across two alignment objectives: harmlessness and honesty. They find that training against probes that do not update during training is easily exploitable, but continuously updated probes reduce harmfulness and improve honesty while preserving utility.

RSS Score 0 10/1/2026, 4:00:00 AM Original Source
Save an API key to vote.