
Thinking Machines Lab open-sources 975B-parameter Inkling, Mixture-of-Experts model
The AMW Read
Novelty=2: new top-tier entrant from high-profile founder changes US open-weight landscape but is a fast-follower in architecture. Significance=2: segment-level shift as Meta abandons open models and Chinese dominance is challenged.
Thinking Machines Lab open-sources 975B-parameter Inkling, Mixture-of-Experts model
Thinking Machines Lab, the startup founded by former OpenAI CTO Mira Murati, has released Inkling, an open-weight Mixture-of-Experts (MoE) transformer model with 975 billion total parameters (41 billion active) and a 1-million-token context window. Trained on 45 trillion tokens of text, image, audio, and video data, the model is distributed under the Apache v2 license via Hugging Face. The company also previewed Inkling-Small, a lightweight variant with 12 billion active parameters. The model performs at roughly one generation behind the latest frontier proprietary models from OpenAI and Anthropic, and trails DeepSeek and GLM on benchmarks, according to the company.
Why it matters: Inkling arrives at a moment when the open-weight model landscape is dominated by Chinese labs (DeepSeek, Qwen, GLM, Kimi), as Meta has ceased releasing large open models after the Llama 4 controversy. Thinking Machines Lab's entry recreates the 'acqui-licensing' pattern — the company raised $2 billion at a $12 billion valuation from investors including Nvidia, AMD, and Cisco, and has assembled a 300-person team drawn heavily from OpenAI, Anthropic, Meta, and Mistral. The open-weight MoE architecture directly copies DeepSeek-V3's design and uses distillation from existing models including Kimi K2.5, which positions Inkling as a fast-follower strategy aimed at commercial fine-tuning and hosting revenue via its Tinker platform.
Grounded expert take: Inkling is not yet competitive with frontier closed models or top Chinese open models on raw benchmarks, but the strategic bet is on Thinking Machines Lab's 'Interactive Model' architecture — a system that allows human-in-the-loop intervention during inference, enabling real-time correction and collaboration across audio, video, and text streams. If this interaction paradigm proves valuable for enterprise customization, Inkling's fine-tunability on Tinker could create a distribution moat that differentiates it from static open-weight releases. The company has acknowledged that future models will be trained entirely on proprietary data, signaling a transition from distillation-driven fast-following to original capability development.
#ThinkingMachinesLab #Inkling #OpenWeight #MixtureOfExperts #MiraMurati #AIInfrastructure



