Model techniques map
Taxonomymodel architecturetoken mixerhybrid layer stacking

taxonomy node · level 3

hybrid layer stacking

13 methods filed at this node or below it, from the sources of 19 models.

model architecture :: token mixer :: hybrid layer stacking

Matching aids for the classifier: hybrid attention architecture; interleaved SWA and global; hybrid Mamba-Transformer; dense attention fallback.

In this branch 13

Everything filed at this node or below it, with one collapsible heading per child node.

filed here 13

Hybrid Attention core · 27 sources · 29 quotes
Hybrid Mamba-Transformer core · 4 sources · 6 quotes
Hybrid Mamba-Attention core · 3 sources · 4 quotes
Gated DeltaNet and Full Attention core · 1 source · 1 quote
Gated DeltaNet and Gated Attention core · 1 source · 2 quotes
Gated DeltaNet and Qwen Sparse Attention core · 1 source · 1 quote
Search-Based Sliding-Window Attention Pattern evaluated · 1 source · 1 quote
Dense Attention Fallback unclear · 1 source · 1 quote

By model

Which of this branch's techniques each model's own documents describe, and how strongly. Under each model: its strongest treatment anywhere in the branch.