Model techniques map
Techniquesmodel architecturecontext capacity

specific method · filed under model architecture

Token normalization for N-gram vocabulary compression

Token normalization used to compress the vocabulary for N-gram embeddings.

source
1
model
1
labs adopt it
0
strongest
evaluated

How sources treat it

One count per evidence span, weakest treatment to strongest.

evaluated 1

Documented in

Evidence

1 span quoted from the sources, strongest treatment first.

including but not limited to token normalization for vocabulary compression ... Despite these efforts, we observed no consistent performance gains in our training recipe.

evaluateddata curationin Qwen3.8-NextQwen

Filed alongside

Other methods under model architecture :: context capacity.