Multi-LLM Collaborative Alignment via Stackelberg Games

A game-theory-inspired framework, Stackelberg Alignment, is proposed for language models (LLMs) to collaborate and improve collectively. The framework uses an EXP3 bandit to select instructions for LLMs to respond to, and the LLMs learn from each other's responses through peer judgment and reputation-based matching.

RSS Score 0 10/1/2026, 4:00:00 AM Original Source
Save an API key to vote.