You changed your mind. The model didn't.
1HKUST · 2Tencent Hy
Demystifying intent in multi-turn dialogue
Language models keep acting on what was said, even after the user rejected or replaced it.
The setup
One task, four conversations
The task arrives turn by turn. After the bag count we insert one event. Only Revised changes the answer.
Figure 1 from the paper

Benchmark
Intent-Eval
414 source tasks from four established benchmarks, each with one controlled revision. Four multi-turn conditions and four single-turn controls give 3,312 evaluation instances per model.
Domain overview
Four domains from established benchmarksEvaluation instances
per modelMulti-turn dialogue
Mean 7.14 turns · range 3–14Domain rules propose edits to arguments, entry points, query conditions or numbers.
Every edit must be unambiguous, atomic and verifiable.
One revision per task, with a new reference answer y(1).
Rendered as one shared proposal: rejected in Retained, accepted in Revised.
Figure 2 from the paper

Findings
Models lose track of what is in effect
Eight models in the main comparison, ten in the error analysis.
Every inserted event lowers accuracy.
Points of multi-turn accuracy lost relative to Original, eight-model mean. Spreading the task over turns already costs 36.30 pp. Even inside one message, a decision costs 5.74 pp (Revised) and 7.52 pp (Retained).
Wrong answers still carry content that is no longer in effect.
Share of new errors (correct in Original, wrong after the decision) that contain inactive content: 446 of 748 and 201 of 649, pooled over ten models.
Retained · GPT-5.6-Luna
Jairus earns $0.80 per task. The user rejects $0.90.
Model answers $8 using $0.90. Correct: $6.
Revised · Claude-Sonnet-5
The car payment changes from 20% to 21%. The user accepts.
Model answers $720 using 20%. Correct: $696.
Losses deepen with repetition and persist after the decision.
Change in accuracy (pp) from each path's k = 0, four-model mean. Four clarifications after a decision still cost 2.72 pp (Retained) and 3.50 pp (Revised).
Figures 4 and 5 from the paper


Method
Intent-OPSD
The frozen Teacher sees the active task; the Student sees the dialogue.
Swipe sideways to see the whole diagram.
Token probabilities and the JS value in the animation are illustrative.
Intent-OPSD lifts multi-turn accuracy.
Mean over four models and four domains. Intent-OPSD beats Base in all 16 model–domain pairs and SFT in 13, and keeps its gains at every added depth (+9.61 to +11.74 pp over Base). Mean Retained accuracy stays close to SFT.
Base vs. SFT vs. Intent-OPSD
Table 2 of the paper. Four models, four domains, four conditions.
Figure 3 from the paper

All results · Table 1
Every model, every domain
Cite
BibTeX
@misc{chen2026changedmindmodeldidnt,
title={You Changed Your Mind, The Model Didn't: Demystifying Intent in Multi-Turn Dialogue},
author={Junle Chen and Wei Chen and Zhengjun Huang and Zhoujin Tian and Yuxuan Liu and Kai Wang and Rui Chen and Xiaofang Zhou},
year={2026},
eprint={2610.06496},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2610.06496},
}