PublishersGemma Team, Google DeepMind
Gemma Team, Google DeepMind
1 document read from this publisher.
Gemma 4 Technical Report
technical report · 2026-07-02 · 29 techniques
Mixture of ExpertsHybrid AttentionMulti-Token PredictionDefault thinking modeQuantization-Aware TrainingEncoder-free multimodal architecturePer-Layer Embeddings (PLE)2D rotary position embeddingEvaluation without safety filtersRaw audio projection into the LLM embedding space2D coordinate-based positional embeddingsAttribution, hedging, and refusal data subsetsData Replica Reduction over the Data Center NetworkFrozen encoders during pre-trainingKey-value reuse in global attention layersKV-cache sharingMobile quantizationPer-Block Scalar ScalingPost-training data filtering for safety and factualityPre-training data filtering for decontamination and safetyQ4_0 quantizationSentencePiece tokenizerSeparate PT and IT end tokensSingle-matrix-multiplication vision projectionSlice-Granularity ElasticityTop-k over token clustersUniversal Speech Model (USM)-based audio encoderVariable aspect-ratio image handlingZeRO-3 Optimizer State Sharding