specific method · filed under model architecture
Token normalization for N-gram vocabulary compression
Token normalization used to compress the vocabulary for N-gram embeddings.
- source
- 1
- model
- 1
- labs adopt it
- 0
- strongest
- evaluated
How sources treat it
One count per evidence span, weakest treatment to strongest.
evaluated 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
including but not limited to token normalization for vocabulary compression ... Despite these efforts, we observed no consistent performance gains in our training recipe.
evaluateddata curationin Qwen3.8-NextQwen
Filed alongside
Other methods under model architecture :: context capacity.