implementation detail · filed under model architecture
Latent-dimension reduction
An implementation approach that shrinks routed-expert input dimensions and reinvests saved capacity in nonlinear budget and expert diversity.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
we shrink the routed expert input dimension to reduce communication and memory costs, and reinvest the saved capacity into increasing the nonlinear budget and expert diversity
usedmodel architecturein NVIDIA Nemotron 3 Super and UltraNVIDIA
Filed alongside
Other methods under model architecture :: channel mixer :: mixture of experts :: latent mixture of experts.