implementation detail · filed under inference & serving
Reasoning-effort resolution and system-prompt injection
A chat template resolves the requested effort to a supported level and inserts that level into the system prompt.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- default
How sources treat it
One count per evidence span, weakest treatment to strongest.
default 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
The chat template resolves effort to max unless reasoning_effort is explicitly "low" or "high" (any other value falls back to max), then injects Reasoning Effort: Low|High|Max into the system prompt.
defaultinference servingin GLM-5.3-FlashZ.ai
Filed alongside
Other methods under inference & serving :: reasoning control.
Configurable reasoning effortChain-of-thought reasoningDefault thinking modeDeployment-time scalar effort controlDisabling reasoning via chat-template configurationInference-time reasoning budget controlInterleaved thinking between tool callsThinking mode selectionCross-turn persistent reasoning historyMaximum thinking effortReasoning parserAlways-on thinking modeclear_thinking chat-template parameterControl-token-enabled thinking modeEffort-dependent exponential token-penalty scheduleGenerate-verify-refine loopMedium-effort reasoning modeParallel-fewest-step samplingQwen3 soft thinking switchTask- and mode-specific sampling parameter recommendationsTest-time compute scalingAdaptive reasoningCapped linear reasoning-token length deductionConfigurable thinking or reasoning mode