implementation detail · filed under inference & serving
Frontier-planning to execution-model routing
Plans are routed to a frontier model while execution is routed to Lightning.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- core
How sources treat it
One count per evidence span, weakest treatment to strongest.
core 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
Plans route up to the frontier, execution routes down to Lightning, ensuring that your tokens are spent efficiently and effectively.
coreinference servingin NeMo SwitchyardNVIDIA
Filed alongside
Other methods under inference & serving :: agentic scaffolding.
Advisor strategyAutomatic tool choicePython toolFunction callingHarmony formatBrowser toolGeneric tool-call parserMCP tool configurationMulti-agent collaborationRole-based instruction hierarchyThinking with toolsVisual Search ToolXML-based tool-call schema with DSML tokenAgent Team modeAgentENVAgentic searchAgentic tool callingApp-server mode with adapted tool schemasAppArmor and eBPF sandbox policiesAssistant output channelsAsynchronous teammate spawningBash commands for context retrievalBash computer-use agentBrowsing tool with domain filtering