specific method · filed under model architecture
Shared Experts Active on Every Token
A configuration in which shared experts are active for every token, in addition to routed experts.
Also called 2 shared experts active on every token.
- sources
- 2
- model
- 1
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
core 2
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
each token is routed to 6 of 256 experts, plus 2 shared experts active on every token
coremodel architecturein InklingThinking Machines Lab
plus 2 shared experts active on every token
coremodel architecturein InklingThinking Machines Lab
Filed alongside
Other methods under model architecture :: channel mixer :: mixture of experts :: shared experts.