implementation detail · filed under inference & serving
Qwen3 soft thinking switch
The /think and /nothink tokens switch between thinking and non-thinking modes.
- sources
- 2
- model
- 1
- labs adopt it
- 0
- strongest
- not used
How sources treat it
One count per evidence span, weakest treatment to strongest.
not used 2
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
Qwen3.5 does not officially support the soft switch of Qwen3, i.e., /thinkand /nothink.
not usedinference servingin Qwen3.5-122B-A10BQwen
Qwen3.5 does not officially support the soft switch of Qwen3, i.e., /think and /nothink.
not usedinference servingin Qwen3.5-397B-A17BQwen
Filed alongside
Other methods under inference & serving :: reasoning control.
Configurable reasoning effortChain-of-thought reasoningDefault thinking modeDeployment-time scalar effort controlDisabling reasoning via chat-template configurationInference-time reasoning budget controlInterleaved thinking between tool callsThinking mode selectionCross-turn persistent reasoning historyMaximum thinking effortReasoning parserAlways-on thinking modeclear_thinking chat-template parameterControl-token-enabled thinking modeEffort-dependent exponential token-penalty scheduleGenerate-verify-refine loopMedium-effort reasoning modeParallel-fewest-step samplingTask- and mode-specific sampling parameter recommendationsTest-time compute scalingAdaptive reasoningCapped linear reasoning-token length deductionConfigurable thinking or reasoning modeControllable thinking effort via system message and per-token cost