PublishersQwen
Qwen
9 documents read from this publisher.
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability
technical report · 2026-08-31 · 81 techniques
Mixture of ExpertsGated DeltaNetMuonGated ResidualManifold-Constrained Hyper-ConnectionsN-gram embedding host-memory offload and prefetchN-gram embedding lookupQwen Sparse AttentionBatch-size warmupUpdated hyperparameter scaling lawCUDA GraphGatedNormSplit fused gradients before orthogonalizationHyper-ConnectionsIndependent Per-Branch NormalizationNesterov momentumPer-Branch Scalar Write GateQK-ClipReuse QSA index selection across speculative decoding stepsActivation or logit clippingAdam without weight decay for the N-gram embedding tableAdamW for attention and GDN output gatesAdamW for gated-residual low-rank projectionsAdamW for input embeddings and output headAdamW for the MoE routerAll-branch Elementwise Read GateAsynchronous Micro-Group pipelineAverage poolingBatch-size warmup with adjusted peak learning rateBatch-size warmup with constant-batch optimum peak learning rateBlock-causal scoringBounded Positive GatesCanzonaContextual gating for N-gram embedding injectionCUDA graph capture of the optimizer stepData-Dependent Residual Read and Write OperatorsDense-attention distillationEight-step Newton–Schulz iterationElevated constant-learning-rate training stability stress testFixed elevated learning rate
Qwen/Qwen3.8-Flash-Next · Hugging Face
first party release · model card · 2026-08-26 · 11 techniques
YaRNDefault thinking modeGated ResidualPreserved thinking history modeN-gram embedding lookupQwen Sparse AttentionMulti-Token PredictionCategory-specific assignment of Muon and AdamWTraining without batch-size warmupGated DeltaNet and Qwen Sparse AttentionIncrease video preprocessing longest-edge limit for higher frame-rate sampling
Qwen/Qwen3.6-27B · Hugging Face
first party release · model card · 2026-04-22 · 7 techniques
Qwen/Qwen3.6-35B-A3B · Hugging Face
first party release · model card · 2026-04-22 · 15 techniques
Mixture of ExpertsYaRNGated DeltaNetPresence PenaltyPrompt-based output standardizationInterleaved thinking between tool callsThinking mode selectionGated AttentionLanguage-model-only serving modeRoPE scalingMCP tool configurationPreserve thinkingDedicated serving enginesOutput length recommendationVideo sampling parameters
Qwen/Qwen3.5-122B-A10B · Hugging Face
first party release · model card · 2026-03-09 · 27 techniques
YaRNDefault thinking modeExcluding prior thinking from conversation historyDisabling reasoning via chat-template configurationEarly fusion multimodal trainingPresence PenaltyPrompt-based output standardizationNEXTN speculative decodingAsynchronous RL frameworks for large-scale agent scaffolds and environment orchestration8 Routed + 1 Shared Experts Activated per TokenGated DeltaNet–sparse MoE hybridMCP tool configurationMulti-step MTP trainingQwen3 soft thinking switchTask-specific sampling parametersAdequate max output length setting (32,768 default; 81,920 for hard benchmarks)Gated DeltaNet with reduced KV-head configurationJSON-structured answer-format promptLongest_edge video-preprocessing configuration for high frame-rate sampling of hour-scale videosMode- and task-specific sampling parameter configurationProgressively Complex Task DistributionsScalable reinforcement learning across million-agent environmentsServing Qwen3.5 with Hugging Face Transformers serverServing Qwen3.5 with KTransformersServing Qwen3.5 with SGLangServing Qwen3.5 with vLLMStep-by-step reasoning with boxed-answer prompt for math benchmarks
Qwen/Qwen3.5-397B-A17B · Hugging Face
first party release · model card · 2026-03-09 · 26 techniques
Multi-Token PredictionYaRNGated DeltaNetDefault thinking modeDiscard-all context managementExcluding prior thinking from conversation historyEarly fusion multimodal trainingPrompt-based output standardizationAsynchronous RL frameworks for large-scale agent scaffolds and environment orchestrationGated AttentionRoPE scalingScalable RL at Agent Scale10 Routed + 1 Shared Experts Activated per TokenContext foldingHybrid Mamba-Transformer Mixture-of-Experts Layer LayoutMulti-step MTP trainingQwen-AgentQwen3 soft thinking switchTensor parallelism (degree 8)adequate output length (32,768 tokens default, 81,920 for hard benchmarks)configurable video frame sampling with fps and do_sample_frameshybrid Gated DeltaNet + sparse MoE architectureQwen3 Coder tool-call parserrecommended sampling parameters for thinking and non-thinking modessetting video preprocessor longest_edge to 469,762,048 for hour-scale video understandingvLLM language-model-only mode
Qwen/Qwen3.5-35B-A3B · Hugging Face
first party release · model card · 2026-03-09 · 24 techniques
Multi-Token PredictionYaRNGated DeltaNetDefault thinking modeExcluding prior thinking from conversation historyDisabling reasoning via chat-template configurationEarly fusion multimodal trainingPresence PenaltyPrompt-based output standardizationNEXTN speculative decodingAsynchronous RL frameworks for large-scale agent scaffolds and environment orchestrationGated AttentionLanguage-model-only serving modeRoPE scaling8 Routed + 1 Shared Experts Activated per TokenContext foldingGated DeltaNet–sparse MoE hybridHybrid Mamba-Transformer Mixture-of-Experts Layer LayoutQwen-AgentTask- and mode-specific sampling parameter recommendationsConfigurable video frame sampling (fps/do_sample_frames) during inferenceScalable RL with progressively complex task distributions in million-agent environmentsTask-appropriate maximum output lengthVideo preprocessor longest-edge configuration
Qwen3.6
first party release · code repo · 5 techniques
Qwen3.8-Flash-Next
first party release · code repo · 17 techniques
Mixture of ExpertsGated DeltaNetMuonGated ResidualN-gram embedding host-memory offload and prefetchN-gram embedding lookupQwen Sparse AttentionCategory-specific assignment of Muon and AdamWUpdated hyperparameter scaling lawSplit fused gradients before orthogonalizationGated DeltaNet and Qwen Sparse Attention hybrid architectureTensor parallelism (degree 4)Compressed lightweight indexerContinuous batchingDynamic Gating of Residual Reads and WritesFour-Branch Residual StreamMuon orthogonalization accuracy refinement