implementation detail · not yet filed
shared-memory caching of preprocessed multimodal inputs
Also called --mm-processor-cache-type shm.
- source
- 1
- model
- 1
- lab adopt it
- 1
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
used 1
Documented in
Evidence
1 span quoted from the sources, strongest treatment first.
For multimodal workloads, use --mm-encoder-tp-mode data for data-parallel vision encoding and --mm-processor-cache-type shm for shared-memory caching of preprocessed multimodal inputs.
usedinference servingin vLLMvLLM