You Changed Your Mind, The Model Didn't

You changed your mind. The model didn't.

Junle Chen1, Wei Chen1,2✉, Zhengjun Huang1, Zhoujin Tian1, Yuxuan Liu1, Kai Wang2, Rui Chen2, Xiaofang Zhou1

1HKUST · 2Tencent Hy

HKUSTTencent Hy

Demystifying intent in multi-turn dialogue

Language models keep acting on what was said, even after the user rejected or replaced it.

The paper as a 90-second story, with narration.

The setup

One task, four conversations

The task arrives turn by turn. After the bag count we insert one event. Only Revised changes the answer.

Figure 1 from the paper
Figure 1: single-turn controls, the multi-turn Original conversation, and the three inserted events.

Benchmark

Intent-Eval

414 source tasks from four established benchmarks, each with one controlled revision. Four multi-turn conditions and four single-turn controls give 3,312 evaluation instances per model.

Domain overview

Four domains from established benchmarks
ActionsBFCL104sourcesFunction callsArguments
CodeHumanEval, LCB100sourcesCode generationEntry points
Word problemsNumerical valuesMathGSM8K103sources
Text-to-SQLQuery conditionsOutput fieldsDatabaseSpider107sources

Evaluation instances

per model
0300600900

Multi-turn dialogue

Mean 7.14 turns · range 3–14
ICandidate enumeration

Domain rules propose edits to arguments, entry points, query conditions or numbers.

IIConstraint filtering

Every edit must be unambiguous, atomic and verifiable.

IIIReference answers

One revision per task, with a new reference answer y(1).

IVReview and rendering

Rendered as one shared proposal: rejected in Retained, accepted in Revised.

Figure 2 from the paper
Figure 2: overview and statistics of Intent-Eval.

Findings

Models lose track of what is in effect

Eight models in the main comparison, ten in the error analysis.

Every inserted event lowers accuracy.

−4.43Neutral · clarification only
−5.55Revised · change accepted
−8.06Retained · change rejected

Points of multi-turn accuracy lost relative to Original, eight-model mean. Spreading the task over turns already costs 36.30 pp. Even inside one message, a decision costs 5.74 pp (Revised) and 7.52 pp (Retained).

Wrong answers still carry content that is no longer in effect.

59.6%Retained · uses the rejected proposal
31.0%Revised · keeps the replaced original

Share of new errors (correct in Original, wrong after the decision) that contain inactive content: 446 of 748 and 201 of 649, pooled over ten models.

Retained · GPT-5.6-Luna

Jairus earns $0.80 per task. The user rejects $0.90.

Model answers $8 using $0.90. Correct: $6.

Revised · Claude-Sonnet-5

The car payment changes from 20% to 21%. The user accepts.

Model answers $720 using 20%. Correct: $696.

Losses deepen with repetition and persist after the decision.

−20.8after four rejected proposals
−8.1after four clarifications

Change in accuracy (pp) from each path's k = 0, four-model mean. Four clarifications after a decision still cost 2.72 pp (Retained) and 3.50 pp (Revised).

Figures 4 and 5 from the paper
Figure 4: inactive content in new Retained and Revised errors by domain.
Figure 5: accuracy change with interaction depth for four models and four domains.

Method

Intent-OPSD

The frozen Teacher sees the active task; the Student sees the dialogue.

v

Swipe sideways to see the whole diagram.

Token probabilities and the JS value in the animation are illustrative.

Intent-OPSD lifts multi-turn accuracy.

+10.81pp over Base
+3.73pp over matched SFT

Mean over four models and four domains. Intent-OPSD beats Base in all 16 model–domain pairs and SFT in 13, and keeps its gains at every added depth (+9.61 to +11.74 pp over Base). Mean Retained accuracy stays close to SFT.

BaseSFTIntent-OPSD

Base vs. SFT vs. Intent-OPSD

Table 2 of the paper. Four models, four domains, four conditions.

Figure 3 from the paper
Figure 3: Intent-OPSD.

All results · Table 1

Every model, every domain

Cite

BibTeX

@misc{chen2026changedmindmodeldidnt,
      title={You Changed Your Mind, The Model Didn't: Demystifying Intent in Multi-Turn Dialogue},
      author={Junle Chen and Wei Chen and Zhengjun Huang and Zhoujin Tian and Yuxuan Liu and Kai Wang and Rui Chen and Xiaofang Zhou},
      year={2026},
      eprint={2610.06496},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2610.06496},
}