implementation detail · filed under data curation
Byte-level Byte Pair Encoding (BPE)
A byte-level BPE tokenizer specified here with a vocabulary size of 250k.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
we use the Qwen3.5 tokenizer (Team, 2026), which adopts byte-level byte-pair encoding with a vocabulary size of 250k (up from 150k)
usedmodel architecturein Qwen3.5-OmniQwen
Filed alongside
Other methods under data curation :: tokenization.