AI startup Thinking Machines, founded by former OpenAI CTO Mira Murati, has unveiled Inkling, an open-weight general-purpose model designed to dramatically cut inference costs. The model uses controllable thinking to let enterprises adjust reasoning depth, consuming only one-third the tokens of rivals like Nemotron 3 Ultra. With 975 billion total parameters in a Mixture-of-Experts architecture, a 1-million-token context window, and native multimodality, Inkling targets the biggest barrier to enterprise AI adoption: cost.
What Happened: An Efficiency-First Open-Weight Model
Inkling is built for efficiency from the ground up. Instead of spending most tokens on processing information, it compresses its thought process, a technique that analyst Bradley Shimmin of Futurum Group calls a shift from “knowing things” to “doing things.” The model supports controllable thinking, allowing enterprises to dial reasoning up or down based on task complexity. Native multimodality—processing images, audio, and video as a continuous whole rather than separately—further boosts efficiency. “That’s a great best practice that’s starting to show up,” Shimmin said. The model was pretrained on 45 trillion tokens of text, images, audio, and video.
Why It Matters: Breaking the Cost Barrier
High inference costs have long hindered enterprise AI deployment. Inkling’s approach—using fewer tokens without sacrificing capability—could lower the total cost of ownership for AI applications. Shimmin noted that by “focusing on the epistemic aspect,” the model becomes more capable at getting things done rather than storing compressed knowledge. XPLAIN AI interprets this as a potential paradigm shift: if efficiency gains are real, mid-sized enterprises that previously found AI too expensive may now enter the market, accelerating adoption across industries.
Our Analysis: Competitive Dynamics and Niche Positioning
Inkling competes not only with OpenAI and Anthropic but also with open-source providers like NVIDIA. However, analyst Lian Jye Su of Omdia doubts enterprises will switch en masse from established models. Instead, Inkling may find a niche in tasks requiring lower compute or limited resources. “There is probably an argument to be made whereby enterprises may prefer to use Thinking Machines’ model when there is a certain task that requires much lower compute,” Su said. XPLAIN AI views Inkling as a complementary tool rather than a universal replacement, at least initially. To differentiate further, Thinking Machines may need to pick a hardware partner or specialize in specific use cases, Su added.
Investment Implications: Winners and Risks
From an investment perspective, the impact is likely medium- to long-term. Potential beneficiaries include enterprises that can cut AI costs, but no specific tickers are confirmed. Potential risks face GPU suppliers like NVIDIA (NVDA), as more efficient models could reduce chip demand for inference. However, Thinking Machines has not yet chosen a hardware partner, and Inkling’s real-world performance remains unverified. XPLAIN AI cautions that the competitive pressure may push all model makers to prioritize efficiency, benefiting the entire ecosystem in the long run.
Counter-Scenarios and Uncertainties
Analysts warn that switching costs and workflow optimization for existing models like Nemotron Ultra will deter many enterprises. Inkling’s open-weight nature also places the burden of tuning and deployment on users, which may slow adoption. Established open-source players like NVIDIA, Meta, and Mistral already dominate, making it tough for a startup to gain traction. The model’s actual performance on reasoning accuracy and hallucination reduction needs independent validation.
Key Metrics to Watch
- Hardware partnerships: If Thinking Machines partners with AMD (AMD) or Intel (INTC) instead of NVIDIA, it could shift the AI chip landscape.
- Enterprise adoption cases: Real-world benchmarks on inference speed, accuracy, and cost savings are critical.
- Competitor responses: How OpenAI, Anthropic, and others adjust their pricing and efficiency strategies will shape the market.
#AI #InferenceCost #ThinkingMachines #Inkling #OpenAI #Anthropic #NVIDIA #EnterpriseAI #EfficiencyRevolution
Sources
- Thinking Machines Rolls Out Broad but Efficient Model — aibusiness · News coverage · Thu, 16 Jul 2026 18:35:21 GMT
Written by: XPLAIN AI Editorial Team · Reviewed by: XPLAIN AI Editorial Desk
This content was drafted with AI assistance based on publicly available sources and reviewed under XPLAIN AI's editorial standards.
