Two Clocks in Diffusion MLLMs: When Answers Stabilize Before Rationales Unfold

This paper examines the behavior of answer candidates and rationales in masked diffusion MLLMs, finding that answers can stabilize before rationales unfold. The study analyzes three visual question-answering benchmarks, revealing differences in answer coverage and observation windows. The findings have implications for the development and optimization of MLLMs.

RSS Score 0 10/2/2026, 4:00:00 AM Original Source
Save an API key to vote.