Model techniques map
Taxonomyinference & servingserving parallelism

taxonomy node · level 2

serving parallelism

17 methods filed at this node or below it, from the sources of 15 models.

inference & serving :: serving parallelism

Matching aids for the classifier: serving-time expert parallelism; data-parallel attention; prefill-decode disaggregation; DeepEP; all-to-all communication; low-precision MoE combine.

In this branch 17

Everything filed at this node or below it, with one collapsible heading per child node.

filed here 17

Tensor parallelism (degree 4) default · 2 sources · 2 quotes
DeepEP used · 4 sources · 4 quotes
Prefill-decode disaggregation used · 4 sources · 4 quotes
Attention data parallelism used · 3 sources · 4 quotes
Encoder-Prefill-Decode disaggregation used · 2 sources · 2 quotes
Tensor parallelism (degree 8) used · 2 sources · 2 quotes
Topology-aware NVLink domain placement used · 2 sources · 2 quotes
Data-parallel vision encoding used · 1 source · 1 quote
Tensor parallelism for MoE layers used · 1 source · 1 quote
TensorRT-LLM all-reduce backend used · 1 source · 1 quote
Expert parallelism optional · 1 source · 1 quote
Low-precision MoE combine optional · 1 source · 1 quote
Token migration for balanced expert placement unclear · 1 source · 1 quote

By model

Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.