implementation detail · filed under evaluation
Inline training evaluators
This approach uses benchmarks as evaluators during training.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- optional
How sources treat it
One count per evidence span, weakest treatment to strongest.
optional 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
Benchmarks can also serve as inline training evaluators via `BenchmarkEvaluator`.
optionalevaluation onlyin tinker-cookbookThinking Machines Lab
Filed alongside
Other methods under evaluation :: evaluation harness.
DeepSeek Harness Minimal modeHarborNeMo Evaluator SDKAtomic binary rubric evaluationAutomatic in-flight checkpoint evaluation schedulingBenchmark evaluation frameworkCross-scaffold evaluationIsolated container per rolloutJSON-structured answer-format promptmini-swe-agent harnessNeMo Gym and NeMo Evaluator-based harnessNeMo SkillsPool