specific method · filed under data curation
SentencePiece tokenizer
A SentencePiece tokenizer described as using split digits, preserved whitespace, and byte-level encodings.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
We use the same tokenizer as Gemini Team [2025] that is, a SentencePiece tokenizer [Kudo and Richardson, 2018] with split digits, preserved whitespace, and byte-level encodings.
useddata curationin Gemma 4Google DeepMind
Filed alongside
Other methods under data curation :: tokenization.