Thinking Machines just dropped Inkling-Small, a new AI model packing 276 billion parameters but activating only 12 billion during use, keeping costs low while delivering serious power. This release marks the first time a US-based lab has brought a frontier-class open-weights model to market that competes with Chinese counterparts on both performance and price.
Inkling-Small shines on key agentic automation benchmarks. It scores 80.2% on SWE-Bench Verified, 64.7% on Terminal Bench 2.1, and hits an impressive 95.1% on the math reasoning test AIME 2026. These scores translate to real-world strengths in software engineering, cybersecurity, and research tasks requiring complex, long-term thinking.
Pricing is a major highlight here. On the Tinker API, serverless inference costs $0.30 per million input tokens and $1.20 per million output tokens, roughly half the input cost of OpenAI’s Luna model, while offering a massive 1-million-token context window and native multimodal support features that set it apart from cheaper alternatives like DeepSeek's V4-Flash-0731.
The company adopts a two-model strategy: Inkling handles the heavy-lifting complex tasks, while Inkling-Small manages routine, efficiency-sensitive workflows. This split allows developers to optimize for both capability and cost depending on task demands. Fine-tuning is also accessible at $1.73 per million tokens via the Tinker training API, enhancing customization for varied applications.
This release arrives amid growing competition and innovation in open-weights AI models, signaling renewed momentum from US labs in an area long dominated by Chinese offerings.
This content is informational and does not serve as financial advice.



