Inkling
AIME 2026: 97.1%, HLE w/ Tools: 53.2%, GPQA Diamond: ~90%, competitive with Nemotron 3 Ultra, GLM 5.2, and DeepSeek V4 Pro across text/agentic/multimodal evaluations.
Architectural Profile & Capabilities
AIME 2026: 97.1%, HLE w/ Tools: 53.2%, GPQA Diamond: ~90%, competitive with Nemotron 3 Ultra, GLM 5.2, and DeepSeek V4 Pro across text/agentic/multimodal evaluations. Artificial Analysis Intelligence Index: 41 | Broad-domain generalist work across text, image, and audio inputs — domain adaptation via fine-tuning on Tinker, agentic coding and tool use, multimodal reasoning (charts, documents, audio transcription/analysis), instruction following, forecasting with calibrated confidence, and controllable-effort reasoning (adjust thinking time to balance speed vs depth). First model in a planned family (Inkling-Small with 276B/12B-active in preview). Apache 2.0 licensed open weights on Hugging Face. **Caveats: Thinking Machines explicitly states Inkling is "not the strongest overall model available today, open or closed" — it is designed as a balanced, customizable base rather than a benchmark topper. AA Intelligence Index of 41 trails frontier proprietary models (GPT-5.6 Sol ~59, Opus 4.8 ~57) and leading open-weights peers (Nemotron 3 Ultra 48, GLM-5.2 ~52). Post-training data included outputs from other open-weight models (Kimi K2.5) — fully self-contained post-training planned for next generation. Very new (Jul 15, 2026) with limited independent benchmark verification.**
Sparse Mixture-of-Experts routing: activates only a subset of parameter experts per token, delivering frontier-level intelligence with exceptional efficiency.
Recommended Workloads & Primary Use Cases
Calls to Inkling are routed through strict Zero Data Retention inference channels. Prompts and outputs are never stored, indexed, or monitored by Thinking Machines.
More Models from Thinking Machines
Compare other engines available in this laboratory.