StateTree: Enhancing Long-Term Dialogue Reasoning via Reinforcement Learning
StateTree is a data-driven RL method for long-term dialogue reasoning in large language models. It constructs a tree-structured path-tracing task from dialogues with verifiable ground truth, allowing the model to traverse and compare records across sessions. StateTree generalizes to longer contexts without full-length RL costs and exhibits capabilities such as cross-session retrieval and temporal reasoning.
Save an API key to vote.