Model techniques map
Techniquessoftware implementationinference engine

implementation detail · filed under software implementation

TensorRT-LLM

An open-source library for high-performance, real-time inference optimization of large language models.

Also called TRTLLM, TensorRT-LLM inference optimization.

sources
2
models
2
lab adopt it
1
strongest
used

How sources treat it

One count per evidence span, weakest treatment to strongest.

optional 2used 1

Documented in

Evidence

3 spans quoted from the sources, strongest treatment first.

TensorRT™-LLM is an open source library built to deliver high-performance, real-time inference optimization for large language models

usedinference servingin TensorRT-LLMNVIDIA

VLLM_FLASHINFER_MOE_BACKEND=latency(TRTLLM-Gen)

optionalsoftware implementationin Nemotron 3 UltraNVIDIA

TRT-LLM deployment snippet provided

optionalsoftware implementationin Nemotron 3 UltraNVIDIA

Filed alongside

Other methods under software implementation :: inference engine.