specific method · filed under model architecture
NoPE sparse multi-head latent attention
Sparse multi-head latent attention layers using NoPE, named as one component of a model combining them with linear-attention layers.
Also called NoPE sparse MLA layers.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
core 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
Its 45-layer language model combines KDA linear-attention layers with NoPE sparse MLA layers
coremodel architecturein GLM-5.3-FlashZ.ai
Filed alongside
Other methods under model architecture :: token mixer :: sparse attention.
DeepSeek Sparse AttentionCompressed Sparse AttentionQwen Sparse AttentionSparse attentionHeavily Compressed AttentionGated DeepSeek Sparse AttentionCSA2 Full ModeKV-outer sparse attentionProgressive sequence-length extension for sparse attentionSparse-attention continued pre-training with joint model and indexer optimizationCross-layer KV and index reuse with statically assigned CSA2 modesFixed-budget sparse-attention selectionFrom-scratch sparse attention training without dense warmupJoint backbone and indexer training under sparse attentionNative Sparse AttentionNatively trained sparsityQSA micro-block compression at ratio 4Sequential block processingSparse retrieval over long contextsSparse softmax attentionToken-wise compressionTwo-stage introduction of sparse attentionTwo-stage sparse attention