PublishersThinking Machines Lab
Thinking Machines Lab
5 documents read from this publisher.
thinkingmachines/Inkling-NVFP4 · Hugging Face
first party release · model card · 2026-07-15 · 14 techniques
Mixture of ExpertsHybrid AttentionDiscrete token encoding for audioHierarchical patch encoderJoint multimodal decoding in a shared hidden spaceMulti-turn open-ended external red-teamingRefusal-suppressed variants for capability estimationToken-level expert routingData curationTraining data deduplication and filteringbenchmarking against public frontier modelsdefense-in-depth with layered input/output classifiersNative multimodal encodingPre-release safety evaluation
Inkling: Our Open-Weights Model
first party release · official blog · 2026-07-15 · 19 techniques
Mixture of ExpertsHybrid AttentionEncoder-free multimodal architecturePython toolRelative attentionSFT bootstrapping with synthetic dataAbstention-aware reward for factual QAControllable thinking effort via system message and per-token costHybrid optimization with Muon and AdamLarge-scale asynchronous RLPost-training on diverse domainsReinforcement Learning for Calibration with Proper Scoring RulesRL with Rubric and Claims GradersSafety training to an internal specificationShort convolutions in attention and residual branchesSigmoid-based MoE router with auxiliary-loss-free load balancingTool-set and schema randomizationTraining to resist censorshipWeight decay coupled to learning-rate squared
Inkling Model Card
first party release · model card · 2026-07-15 · 10 techniques
Hybrid AttentionNVFP4Discrete token encoding for audioHierarchical patch encoderMulti-turn open-ended external red-teamingRefusal-suppressed variants for capability estimationData curationShared Experts Active on Every TokenDefense-in-depth for safetyInput/output classification with moderation tools
thinkingmachines/Inkling · Hugging Face
first party release · model card · 2026-07-15 · 15 techniques
Mixture of ExpertsHybrid AttentionDiscrete token encoding for audioHierarchical patch encoderJoint multimodal decoding in a shared hidden spaceMulti-turn open-ended external red-teamingRefusal-suppressed variants for capability estimationShared Experts Active on Every TokenTraining data deduplication and filteringApplication-layer safeguards (content filtering, rate limiting, monitoring)Dangerous-capability uplift assessmentHuman oversight and review for high-stakes model outputsInput/output moderation classifiers layered around the model (defense-in-depth)Synthetic data generation and augmentationTop-6 expert routing
Latest stable release from PyPI
software documentation · code repo · 6 techniques