specific method · filed under inference & serving
Cross-turn persistent reasoning history
Prior reasoning blocks remain available in the conversation context across turns, including tool-calling turns.
Also called persistent reasoning-history chat template, persistent thinking history.
- sources
- 3
- models
- 4
- labs adopt it
- 3
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 3
Documented in
Evidence
3 spans quoted from the sources, strongest treatment first.
It includes integrated thinking mode with reasoning traces preserved across multi-turn conversations
usedotherin Qwen3.6-35B-A3BAlibaba Cloud
all reasoning content is fully preserved throughout the entire conversation.
usedpost trainingin DeepSeek-V4DeepSeek
We train with persistent thinking history, where all prior reasoning blocks remain visible in the context during each generation step.
usedsoftware implementationin Laguna XS.2Poolside
Filed alongside
Other methods under inference & serving :: reasoning control.
Configurable reasoning effortChain-of-thought reasoningDefault thinking modeDeployment-time scalar effort controlDisabling reasoning via chat-template configurationInference-time reasoning budget controlInterleaved thinking between tool callsThinking mode selectionMaximum thinking effortReasoning parserAlways-on thinking modeclear_thinking chat-template parameterControl-token-enabled thinking modeEffort-dependent exponential token-penalty scheduleGenerate-verify-refine loopMedium-effort reasoning modeParallel-fewest-step samplingQwen3 soft thinking switchTask- and mode-specific sampling parameter recommendationsTest-time compute scalingAdaptive reasoningCapped linear reasoning-token length deductionConfigurable thinking or reasoning modeControllable thinking effort via system message and per-token cost