implementation detail · filed under software implementation
FlashInfer
A kernel backend/library referenced as a backend choice; the evidence does not establish a more specific mechanism.
Also called FlashInfer backend.
- sources
- 3
- models
- 2
- labs adopt it
- 2
- strongest
- used
How sources treat it
One count per evidence span, weakest treatment to strongest.
optional 2used 1
Documented in
Further reading
Picked by hand, not extracted: where to read more, not evidence for anything on this page.
- FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving (Ye et al., 2025) paper arxiv.orgintroduces FlashInfer
- flashinfer-ai/flashinfer code github.comofficial kernel library repository
Evidence
3 spans quoted from the sources, strongest treatment first.
mamba-backend flashinfer
usedsoftware implementationin Nemotron 3 UltraNVIDIA
--mamba-backend flashinfer
optionalsoftware implementationin Nemotron 3 UltraNVIDIA
--mm-attention-backend fa4
optionalsoftware implementationin Step 3.7 FlashStepFun
Filed alongside
Other methods under software implementation :: kernel & quantization library.