specific method · filed under inference & serving
Maximum thinking effort
The model is run at the maximum thinking-effort setting; evidence notes that max mode may let the model set its own test-time compute budget.
Also called Max thinking effort, max thinking.
- sources
- 3
- models
- 3
- labs adopt it
- 3
- strongest
- default
How sources treat it
One count per evidence span, weakest treatment to strongest.
optional 1default 2
Documented in
Evidence
3 spans quoted from the sources, strongest treatment first.
In max mode the model sets its own test-time compute budget.
defaultinference servingin Laguna S 2.1Poolside
At launch, Kimi K3 will use max thinking effort by default, with low- and high-effort modes to be introduced in subsequent updates.
defaultinference servingin Kimi K3Moonshot AI
the recommendation of always using the model on Max thinking effort
optionalinference servingin GLM-5.2Z.ai
Filed alongside
Other methods under inference & serving :: reasoning control.
Configurable reasoning effortChain-of-thought reasoningDefault thinking modeDeployment-time scalar effort controlDisabling reasoning via chat-template configurationInference-time reasoning budget controlInterleaved thinking between tool callsThinking mode selectionCross-turn persistent reasoning historyReasoning parserAlways-on thinking modeclear_thinking chat-template parameterControl-token-enabled thinking modeEffort-dependent exponential token-penalty scheduleGenerate-verify-refine loopMedium-effort reasoning modeParallel-fewest-step samplingQwen3 soft thinking switchTask- and mode-specific sampling parameter recommendationsTest-time compute scalingAdaptive reasoningCapped linear reasoning-token length deductionConfigurable thinking or reasoning modeControllable thinking effort via system message and per-token cost