taxonomy node · level 2
tokenization
8 methods filed at this node or below it, from the sources of 8 models.
data curation :: tokenization
Matching aids for the classifier: tokenizer; special tokens; quick instruction tokens.
In this branch 8
Everything filed at this node or below it, with one collapsible heading per child node.
By model
Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.
| Model | techniques |
|---|---|
| GLM-5.3-Flash used | Chat-template image/video placeholder expansion used— |
| DeepSeek-V4-Flash used | Quick Instruction tokens used— |
| DeepSeek-V4-Pro used | Quick Instruction tokens used— |
| Gemma 4 31B used | SentencePiece tokenizer usedSeparate PT and IT end tokens used— |
| Kimi K3 core | XTML (eXtensible Token Markup Language) core— |
| Laguna-S-2.1 used | Subtoken averaging used— |
| Qwen3.5-397B-A17B used | Byte-level Byte Pair Encoding (BPE) used— |
| gpt-oss-120b used | o200k_harmony tokenizer used— |