implementation detail · filed under data curation
Joint prefetching and assignment of training samples
Jointly prefetches and assigns training samples to minimize overlap across pre-training and context extension.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 2
Documented in
Evidence
2 spans quoted from the sources, strongest treatment first.
We minimize sample overlap during pre-training and context extension by jointly prefetching and assigning training samples.
usedunclearin DeepSeek-V4.1-FlashDeepSeek
We minimize sample overlap during pre-training and context extension by jointly prefetching and assigning training samples
usedunclearin DeepSeek-V4.1-FlashDeepSeek
Filed alongside
Other methods under data curation :: deduplication.