Character Training for Risk-Averse Agents

Researchers explored character training as a method to instill risk aversion in AI agents to prevent misaligned behavior. They trained models using constant absolute risk aversion (CARA) and found that character-trained models performed competitively with baselines and better out-of-distribution. Token budget and model choice are key factors in instilling risk aversion.

RSS Score 0 9/30/2026, 4:00:00 AM Original Source
Save an API key to vote.