Model techniques map
Techniquesmodel architecturechannel mixermixture of expertslatent mixture of experts

implementation detail · filed under model architecture

Latent-dimension reduction

An implementation approach that shrinks routed-expert input dimensions and reinvests saved capacity in nonlinear budget and expert diversity.

source
1
model
1
lab adopt it
1
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

used 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

we shrink the routed expert input dimension to reduce communication and memory costs, and reinvest the saved capacity into increasing the nonlinear budget and expert diversity

usedmodel architecturein NVIDIA Nemotron 3 Super and UltraNVIDIA

Filed alongside

Other methods under model architecture :: channel mixer :: mixture of experts :: latent mixture of experts.