specific method · filed under inference & serving
Always-on thinking mode
Reasoning is enabled unconditionally in generation, rather than selected or disabled per request.
Also called always-on thinking block, always has thinking enabled.
- sources
- 2
- models
- 2
- labs adopt it
- 2
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
core 2
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
Kimi K3 always has thinking enabled, and will return `reasoning_content`.
coreinference servingin Kimi K3Moonshot AI
Thinking is always on — the generation prompt opens a thinking block unconditionally.
coreinference servingin GLM-5.3-FlashZ.ai
Filed alongside
Other methods under inference & serving :: reasoning control.
Configurable reasoning effortChain-of-thought reasoningDefault thinking modeDeployment-time scalar effort controlDisabling reasoning via chat-template configurationInference-time reasoning budget controlInterleaved thinking between tool callsThinking mode selectionCross-turn persistent reasoning historyMaximum thinking effortReasoning parserclear_thinking chat-template parameterControl-token-enabled thinking modeEffort-dependent exponential token-penalty scheduleGenerate-verify-refine loopMedium-effort reasoning modeParallel-fewest-step samplingQwen3 soft thinking switchTask- and mode-specific sampling parameter recommendationsTest-time compute scalingAdaptive reasoningCapped linear reasoning-token length deductionConfigurable thinking or reasoning modeControllable thinking effort via system message and per-token cost