摘要 · Summary
AI · Source-Checked BriefTechOrange (廖紹伶) · 16 Jul 2026
Executive summary · one-page brief
The Open Model That Admits It's Not the Strongest
Thinking Machines Lab — founded by ex-OpenAI CTO Mira Murati — released its first open-weight model Inkling on 15 Jul 2026: a 975B / 41B-active multimodal MoE with 1M-token context, Apache 2.0, a thinking-effort dial, and direct answers on politically censored topics with safety scores intact. TechOrange frames it as the Western answer to Chinese open-source dominance — and cross-checked against the official post and ten outlets, the report is faithful. Our audit adds refinements: GPT-5.5 (not "GPT-5"), the ~1/13 cost holds only vs Claude Opus, the Bridgewater case is a vendor co-publication — and Inkling itself partly learned from Kimi K2.5.
By the Numbers all figures verified vs official post
- 975B/41Bmultimodal MoE — total / active parameters
- 1Mmax context on open weights (served tiers 256K / 64K)
- 0.2–0.99controllable thinking-effort dial
- 98.6%StrongREJECT — safety intact despite censorship non-compliance
What Follows 起 · 承 · 轉 · 合
| Beat | Page | What it covers |
|---|---|---|
| 起 · 承 The model | 2 | Specs, Apache 2.0, the effort dial, CoT condensation, the censorship stance. |
| 轉 The bet & the audit | 3 | Adoption economics via Tinker; the Bridgewater case with caveats; figure audit. |
| 合 The frame & its irony | 4 | The West-vs-China framing — and the Kimi K2.5 distillation TC omits. |
Bottom line
TechOrange's synthesis is faithful — no fabricated figures. Every defect found is a refinement (naming, scope, sourcing), not a debunking.
Reframe
The sharpest omitted fact: the "Western answer" to Chinese open models was itself partly distilled from Kimi K2.5.
起承轉合 · The Model, the Bet, the Frame · 1 / 1
起 — What Inkling Is released 15 Jul 2026 (US time)
Licence · availability
Open Weights, Apache 2.0
- Weights on Hugging Face (thinkingmachines/Inkling) — download, run, customise; no royalties.
- Runs on SGLang, vLLM, TokenSpeed, llama.cpp; NVFP4 build for Blackwell (≥600 GB).
Architecture · scale
975B Multimodal MoE
- Natively multimodal — text, image, audio reasoning; 975B total / 41B active.
- 1M-token context on the open weights; Tinker/API tiers serve 256K + 64K.
承 · 轉 — Three Highlights, and the Business Bet dial · condensation · adoption economics
A thinking-effort dial sweeps 0.2–0.99, trading depth for latency. Past 30M RL rollouts the chain of thought spontaneously condensed — "more concise over time, dropping grammatical overhead while remaining comprehensible" — cutting latency sharply. Cognition's Censorship Eval shows "strong patterns of censorship non-compliance" — direct answers where others refuse — while StrongREJECT 98.6% and FORTRESS 78.0% / 95.9% hold. TM concedes open weights allow refusal removal, advising defence-in-depth plus Llama Guard.
- 0.2–0.99thinking-effort dial — depth vs latency
- 30M+RL rollouts → spontaneous CoT condensation
TC's core read: Inkling bets on enterprise-adoption economics, not leaderboard rank — monetised through Tinker fine-tuning and managed services, shifting spend from renting closed APIs to owned, customised infrastructure. The showcase: Bridgewater fine-tuned Qwen3-235B on Tinker for financial-document classification — 84.7% vs Claude Opus 78.2% (~$100/1k tasks), GPT-5.5 78.0% (~$70), at ~$7.25/1k. But the result is a joint Bridgewater/TM co-publication — vendor-interested, no independent replication — and one nuance TC softens: Tinker itself is token-metered; the real contrast is owned-custom vs rented-closed-API, not metered vs unmetered.
- 84.7%fine-tuned Qwen3-235B — vs Opus 78.2 / GPT-5.5 78.0
- 13.8×cost edge — vs Claude Opus only; ~9.7× vs GPT-5.5
Figure Audit refinements, not debunking
| Item | TechOrange | Audit finding |
|---|---|---|
| Model name | "GPT-5" in the Bridgewater comparison | Correct name: GPT-5.5 (78.0%, ~$70/1k tasks) |
| ~1/13 cost | paired with both GPT-5(.5) and Claude Opus | holds vs Opus only (13.8×); ~9.7× vs GPT-5.5 |
| Inkling-Small | 「2,7?0 億」 — digits cut; active count omitted | 276B total / 12B active, weights pending |
合 — The West-vs-China Frame & Its Irony and what it leaves out
TC asks whether open AI is still "only a Chinese story": with Llama 4 underperforming and Meta closing up, enterprises moved to Qwen, DeepSeek, GLM and Kimi — Inkling is the Western counter-move, with TM's manifesto casting the closed frontier paradigm as "central planning" and citing Hayek on dispersed knowledge. All confirmed. What TC omits is the twist: Inkling's early post-training data was partly distilled from Kimi K2.5 — the Western answer partly learned from the Chinese models it answers. Also unmentioned: day-0 Baseten/Databricks support, and that future TM models are case-by-case — not necessarily open. Where it lands: open-vs-closed is undecided, but the West has pushed a new chip onto the table (「西方陣營至少已經把一張新的籌碼推上了牌桌」) — the chip is real; read its fine print.
- K2.5Kimi K2.5 — the distillation source TC omits
- 45Ttokens of multimodal pretraining (also omitted)