PublishersLMSYS
LMSYS
1 document read from this publisher.
Qwen3.8-Flash-Next: Day-0 Support in SGLang
technical report · vendor docs · 2026-08-26 · 18 techniques
Gated ResidualN-gram embedding host-memory offload and prefetchN-gram embedding lookupQwen Sparse AttentionGated DeltaNet and Qwen Sparse Attention hybrid architectureReuse QSA index selection across speculative decoding stepsFour-Stream Hyper-Connection Combine UpdateFusion of activation, gating, and reduction into GEMM epiloguesHidden-dimension split Combine kernelIndexShare for multi-token prediction speculative decodingLow-Rank Gated MixOffline weight reordering for local gate reductionOverlap pinned-host embedding gather with decoder computationQSA micro-block compression at ratio 4Shape-aware kernel dispatch for Hyper-ConnectionSparse pinned-host offloadSplit-K CuTe GEMM for low-batch Hyper-Connection MixTop-512 block selection