implementation detail · filed under model architecture
Full Attention Residuals
An attention-residual variant using a layer-specific learnable pseudo-query to retrieve representations.
- source
- 1
- model
- 1
- labs adopt it
- 0
- strongest
- mentioned
How sources treat it
One count per evidence span, weakest treatment to strongest.
mentioned 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
For each layer l, we define a layer-specific learnable pseudo-query ql =w l ∈R d and keys and values
mentionedmodel architecturein Kimi K3Moonshot AI
Filed alongside
Other methods under model architecture :: normalization & residual.
Gated ResidualManifold-Constrained Hyper-ConnectionsAttention Residuals (AttnRes)GatedNormIdentity Hyper-ConnectionsRMSNormSingle-Pass mHCBlock AttnResHyper-ConnectionsIndependent Per-Branch NormalizationPer-Branch Scalar Write GateQK-ClipAll-branch Elementwise Read GateBounded Positive GatesData-Dependent Residual Read and Write OperatorsDynamic Gating of Residual Reads and WritesElementwise Data-Dependent Residual Read GateFour-Branch Residual StreamFour-Stream Hyper-Connection Combine UpdateLow-Rank Gated MixPost-Embedding RMSNormPre-LNSimplified AltUpSparse Gated Residual Writes