specific method · filed under inference & serving
Modality-specific deployment
Reducing memory footprint by deploying only the modalities needed for a use case.
Also called deploying only the modalities you need.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- optional
How sources treat it
One count per evidence span, weakest treatment to strongest.
optional 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
Because our audio and vision encoders are not needed in many use cases, you can optimize your memory footprint even further by deploying only the modalities you need.
optionalinference servingin Gemma 4 E2B text-only modelGoogle
Filed alongside
Other methods under inference & serving :: context management.
Discard-all context managementPreserved thinking history modeExcluding prior thinking from conversation historyContext compactionContext foldingDiscard-75%Preserve thinkingThinking context management for tool useTrajectory summarization and rollout re-initiationContext management methodContext management strategyDiscarding tool-call historyHierarchical context managementKeep-recent-kMemory compressionSummary-based context compressionTest-time context management for extending token budgets