implementation detail · filed under inference & serving
clear_thinking chat-template parameter
A chat-template flag for controlling thinking content, documented with a default of false.
Also called clear_thinking chat template flag, clear_thinking, clear_thinking parameter in chat template.
- sources
- 2
- models
- 2
- lab adopt it
- 1
- strongest
- default
How sources treat it
One count per evidence span, weakest treatment to strongest.
optional 1default 1
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
In the chat template for GLM-5.3, clear_thinking defaults to false if not passed. For chat scenarios, explicitly pass clear_thinking=true.
defaultunclearin GLM-5.3Z.ai
In the chat template for GLM-5.3-Flash, clear_thinking defaults to false if not passed.
optionalinference servingin GLM-5.3-FlashZ.ai
Filed alongside
Other methods under inference & serving :: reasoning control.
Configurable reasoning effortChain-of-thought reasoningDefault thinking modeDeployment-time scalar effort controlDisabling reasoning via chat-template configurationInference-time reasoning budget controlInterleaved thinking between tool callsThinking mode selectionCross-turn persistent reasoning historyMaximum thinking effortReasoning parserAlways-on thinking modeControl-token-enabled thinking modeEffort-dependent exponential token-penalty scheduleGenerate-verify-refine loopMedium-effort reasoning modeParallel-fewest-step samplingQwen3 soft thinking switchTask- and mode-specific sampling parameter recommendationsTest-time compute scalingAdaptive reasoningCapped linear reasoning-token length deductionConfigurable thinking or reasoning modeControllable thinking effort via system message and per-token cost