implementation detail · filed under inference & serving
Control-token-enabled thinking mode
Thinking is enabled by adding the <|think|> token to the system prompt and disabled by removing it.
Also called configurable thinking modes.
- sources
- 2
- model
- 1
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
optional 1core 1
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
Thinking is enabled by including the <|think|>token at the start of the system prompt. To disable thinking, remove the token.
coreinference servingin Gemma 4Google DeepMind
Thinking is enabled by including the <|think|>token at the start of the system prompt. To disable thinking, remove the token.
optionalinference servingin Gemma 4Google DeepMind
Filed alongside
Other methods under inference & serving :: reasoning control.
Configurable reasoning effortChain-of-thought reasoningDefault thinking modeDeployment-time scalar effort controlDisabling reasoning via chat-template configurationInference-time reasoning budget controlInterleaved thinking between tool callsThinking mode selectionCross-turn persistent reasoning historyMaximum thinking effortReasoning parserAlways-on thinking modeclear_thinking chat-template parameterEffort-dependent exponential token-penalty scheduleGenerate-verify-refine loopMedium-effort reasoning modeParallel-fewest-step samplingQwen3 soft thinking switchTask- and mode-specific sampling parameter recommendationsTest-time compute scalingAdaptive reasoningCapped linear reasoning-token length deductionConfigurable thinking or reasoning modeControllable thinking effort via system message and per-token cost