specific method · filed under training objective
Dense-attention distillation
Distills the backbone’s full-sequence attention distribution into an indexer.
Also called dense distillation.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
Stage 1: Dense Distillation. We first distill the full-sequence attention distribution of the backbone into the indexer.
usedunclearin Qwen3.8-Flash-NextQwen
Filed alongside
Other methods under training objective :: distillation objective.