implementation detail · filed under model architecture
CSA2 Full Mode
One of CSA2's statically assigned layer modes; the evidence distinguishes it from Reindex and Reuse but does not further specify its operation.
Also called Full Mode.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
optional 1used 1
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
In each group, the first layer operates in Full Mode, and the remaining five layers operate in Reuse Mode.
usedunclearin DeepSeek-V4.1-FlashDeepSeek
Each CSA2 layer is statically assigned one of three modes: Full, Reindex, or Reuse.
optionalinference servingin DeepSeek-V4.1-FlashDeepSeek
Filed alongside
Other methods under model architecture :: token mixer :: sparse attention.
DeepSeek Sparse AttentionCompressed Sparse AttentionQwen Sparse AttentionSparse attentionHeavily Compressed AttentionGated DeepSeek Sparse AttentionKV-outer sparse attentionProgressive sequence-length extension for sparse attentionSparse-attention continued pre-training with joint model and indexer optimizationCross-layer KV and index reuse with statically assigned CSA2 modesFixed-budget sparse-attention selectionFrom-scratch sparse attention training without dense warmupJoint backbone and indexer training under sparse attentionNative Sparse AttentionNatively trained sparsityNoPE sparse multi-head latent attentionQSA micro-block compression at ratio 4Sequential block processingSparse retrieval over long contextsSparse softmax attentionToken-wise compressionTwo-stage introduction of sparse attentionTwo-stage sparse attention