PublishersKili Technology
Kili Technology
1 document read from this publisher.
DeepSeek V4 Guide: Engram Memory, Training Data Strategy & Release Status (2026)
news analysis · third party analysis · 2026-03-23 · 33 techniques
Mixture of ExpertsConfigurable reasoning effortDeepSeek Sparse AttentionGroup Relative Policy OptimizationSupervised fine-tuningMuonReinforcement LearningOn-Policy DistillationManifold-Constrained Hyper-ConnectionsGenerative Reward ModelCompressed Sparse AttentionEngramHeavily Compressed AttentionBatch-invariant deterministic kernelsFP4 Quantization-Aware TrainingSample-level attention maskingSwiGLU clampingAnticipatory RoutingFiltering batched auto-generated and templated contentFull-vocabulary logit distillationInternal engineering task evaluationQuick Instruction tokensXML-based tool-call schema with DSML tokenLong-document curationMid-training phaseMultilingual corpus expansionPairwise comparison for Chinese writingPreemptible rollout serviceRubric-guided RL dataSandbox infrastructure (DSec)Specialist-model trainingTask-specific data injection during mid-trainingWhite-collar enterprise task evaluation