PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization
Researchers propose Prefix-Aware Internal Reward (PAIR), a two-stage model for LLMs to address limitations in credit assignment across intermediate steps in complex tasks.
Save an API key to vote.