specific method · filed under model architecture
Sparse-attention continued pre-training with joint model and indexer optimization
A continued-pretraining stage that adapts model parameters to a sparse pattern after indexer warm-up.
Also called Sparse Training Stage.
- sources
- 2
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 2
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
Following indexer warm-up, we introduce the fine-grained token selection mechanism and optimize all model parameters to adapt the model to the sparse pattern of DSA.
usedoptimizationin DeepSeek-V3.2DeepSeek
Following indexer warm-up, we introduce the fine-grained token selection mechanism and optimize all model parameters to adapt the model to the sparse pattern of DSA.
usedoptimizationin DeepSeek-V3.2DeepSeek
Filed alongside
Other methods under model architecture :: token mixer :: sparse attention.
DeepSeek Sparse AttentionCompressed Sparse AttentionQwen Sparse AttentionSparse attentionHeavily Compressed AttentionGated DeepSeek Sparse AttentionCSA2 Full ModeKV-outer sparse attentionProgressive sequence-length extension for sparse attentionCross-layer KV and index reuse with statically assigned CSA2 modesFixed-budget sparse-attention selectionFrom-scratch sparse attention training without dense warmupJoint backbone and indexer training under sparse attentionNative Sparse AttentionNatively trained sparsityNoPE sparse multi-head latent attentionQSA micro-block compression at ratio 4Sequential block processingSparse retrieval over long contextsSparse softmax attentionToken-wise compressionTwo-stage introduction of sparse attentionTwo-stage sparse attention