Model techniques map
Techniquesmodel architecturetoken mixersparse attention

specific method · filed under model architecture

Joint backbone and indexer training under sparse attention

Training that jointly updates the model backbone and indexer after sparse attention is enabled.

Also called sparse training.

source
1
model
1
lab adopt it
1
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

During the final stage of CPT, QSA is enabled and the backbone and indexer are jointly trained for 8,000 steps

usedunclearin Qwen3.8-Flash-NextQwen

Filed alongside

Other methods under model architecture :: token mixer :: sparse attention.