specific method · filed under data curation
o200k_harmony tokenizer
A tokenizer using the o200k_harmony token vocabulary.
Also called o200k_harmony Byte Pair Encoding tokenizer.
- sources
- 2
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 2
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
Across all training stages, we utilize our o200k_harmony tokenizer
usedmodel architecturein gpt-oss-120b and gpt-oss-20bOpenAI
They are part of a new token vocabulary called o200k_harmony, which landed in OpenAI’s tiktoken tokenizer library this morning.
usedsoftware implementationin openai/harmonyOpenAI
Filed alongside
Other methods under data curation :: tokenization.