StateTree: Enhancing Long-Term Dialogue Reasoning via Reinforcement Learning

StateTree is a data-driven RL method for long-term dialogue reasoning in large language models. It constructs a tree-structured path-tracing task from dialogues with verifiable ground truth, allowing the model to traverse and compare records across sessions. StateTree generalizes to longer contexts without full-length RL costs and exhibits capabilities such as cross-session retrieval and temporal reasoning.

RSS Score 0 10/1/2026, 4:00:00 AM Original Source
Save an API key to vote.