Model techniques map

28 open-weight models · 164 source documents · every claim quoted

What’s actually inside today’s open‑weight models.

We read the technical reports, model cards and repos behind the most-used open models and pulled out every method they describe — each one linked to the exact sentence that says so.

The consensus recipe

If you built a model from what the field agrees on, it would look like this.

The methods the most labs report adopting at each stage of the pipeline, one per sub-area. Numbers count labs, not models, so a lab that documents seven variants in one report still counts once — 13 labs in all.

  1. Data curation

  2. Training objective

  3. Model architecture

  4. Optimization

  5. Post-training

  6. Inference & serving

Bar length is relative to Mixture of Experts, adopted by 13 of 13 labs. Adopted means a document states the model uses it, ships it as an option, or builds on it — not only that it discusses it.

Who uses what

One grid, 13 labs, the methods where they part ways.

Each column is one lab’s most thoroughly documented model. Rows are methods common enough to look for in a model’s documentation, picked where the labs’ answers differ most — the ones every lab adopts are in the recipe above.

A blank cell means the model’s documents don’t say — not that the model doesn’t use it.

A ✓ beside a cell means the model’s own source code shows the method, whatever its documents say; a ✕ that the code does not have it.

MethodDeepSeek-V4.1-FlashNVIDIA-Nemotron-3-Ultra-550B-A55BHy3GLM-5.2MiniMax-M3Gemma 4 31BInklingKimi K3Laguna-S-2.1MiMo-V2.6-ProQwen3.8-Flash-NextStep-3.7-Flashgpt-oss-120b
model architecture
Hybrid Attentionnot in its documentsnot in its documentsnot in its documentsnot in its documentsnot in its documentscorecorecorecorecorecorenot in its documentsused
Multi-Token Predictionnot in its documents; in its codecore; in its codecore; not in its codeusednot in its documents; in its codenot in its documents; not in its codenot in its documents; in its codenot in its documents; not in its codenot in its documents; not in its codecorecore; in its codeoptional; in its codenot in its documents; not in its code
Grouped-query attentionnot in its documents; not in its codenot in its documents; in its codecore; in its codenot in its documents; not in its codecore; in its codenot in its documents; in its codeused; in its codenot in its documents; not in its codecore; in its codenot in its documents; in its codenot in its documents; in its codenot in its documents; in its codeused; in its code
optimization
Muonusednot in its documentsnot in its documentsusednot in its documentsnot in its documentsnot in its documentsusedcorenot in its documentscorenot in its documentsnot in its documents
Expert Parallelismnot in its documentsusednot in its documentsusednot in its documentsnot in its documentsnot in its documentsusednot in its documentsnot in its documentsnot in its documentsusednot in its documents
Warmup-Stable-Decaynot in its documentsusednot in its documentsnot in its documentsnot in its documentsnot in its documentsnot in its documentsnot usedusednot in its documentsnot in its documentsnot in its documentsnot in its documents
post-training
Asynchronous reinforcement learningnot in its documentsusednot in its documentsusednot in its documentsnot in its documentsnot in its documentsnot in its documentsnot in its documentscorenot in its documentsnot in its documentsnot in its documents
inference & serving
Speculative decodinguseddefaultcoreusednot in its documentscorenot in its documentsnot in its documentsnot in its documentsnot in its documentsnot in its documentsoptionalnot in its documents
FP8 KV-cache quantizationnot in its documentsusedoptionalnot in its documentsnot in its documentsnot in its documentsnot in its documentsnot in its documentsusednot in its documentsnot in its documentsusednot in its documents
CUDA Graphnot in its documentsnot in its documentsnot in its documentsnot in its documentsusednot in its documentsnot in its documentsnot in its documentsnot in its documentsnot in its documentsusednot in its documentsused
coredefaultusedoptionalnot usednot in its documents✓in its code✕not in its code
Every method by model, area by area →

Most widely attested

Ranked by how many documents describe the method, not by how often a single document repeats it.

TechniqueAreaSourcesQuotes
Mixture of Expertsmodel architecture88101
Multi-Token Predictionmodel architecture3040
Configurable reasoning effortinference & serving3334
Hybrid Attentionmodel architecture2729
Speculative decodinginference & serving2225
Multi-Token Predictioninference & serving2020
YaRNmodel architecture1219
DeepSeek Sparse Attentionmodel architecture1517
Group Relative Policy Optimizationpost-training1517
Multi-Teacher On-Policy Distillationpost-training1217
Supervised fine-tuningpost-training1417
FP8 KV-cache quantizationinference & serving1315

All 1733 techniques →

Areas

The nine roots of the curated method taxonomy. A method is filed into a path that already exists; what fits nowhere becomes a proposal, listed on About.

Models covered

RankModelCreatorTechniquesSources
1GLM-5.3-FlashZ.ai (Zhipu AI)335
2DeepSeek-V4.1-FlashDeepSeek2014
3Hy4-previewTencent (Hunyuan)377
4DeepSeek-V4-Flash-0731DeepSeek174
5NVIDIA-Nemotron-3-Ultra-550B-A55BNVIDIA2087
6MiMo-V2.6-FlashXiaomi1365
7DeepSeek-V4-FlashDeepSeek824
8MiMo-V2.5Xiaomi778
9GLM-5.3Z.ai (Zhipu AI)466
10Hy3Tencent (Hunyuan)387
11GLM-5.2Z.ai (Zhipu AI)725
12MiniMax-M3MiniMax837
—DeepSeek-V3.2DeepSeek684
—DeepSeek-V4-Flash-Vision-ExpDeepSeek253
—DeepSeek-V4-ProDeepSeek985
—DeepSeek-V4-Pro-0813DeepSeek133
—Gemma 4 31BGoogle DeepMind7810
—InklingThinking Machines Lab7311
—Kimi K3Moonshot AI18310
—Laguna-S-2.1Poolside1638
—MiMo-V2.5-ProXiaomi416
—MiMo-V2.6-ProXiaomi1435
—NVIDIA-Nemotron-3.5-Lightning-30B-A3BNVIDIA737
—Qwen3.5-397B-A17BAlibaba (Qwen)988
—Qwen3.6-35B-A3BAlibaba (Qwen)466
—Qwen3.8-Flash-NextAlibaba (Qwen)1267
—Step-3.7-FlashStepFun426
—gpt-oss-120bOpenAI866

Rank is the model's position among open-weight models in OpenRouter's token ranking for the week of 2026-09-21 (a partial week), the same week for every model; — means outside that week's top 12.

Extraction last run 2026-09-25.