中文

AI · Source-Checked BriefTechOrange (廖紹伶) · 16 Jul 2026

Executive summary · one-page brief

The Open Model That Admits It's Not the Strongest

Thinking Machines Lab — founded by ex-OpenAI CTO Mira Murati — released its first open-weight model Inkling on 15 Jul 2026: a 975B / 41B-active multimodal MoE with 1M-token context, Apache 2.0, a thinking-effort dial, and direct answers on politically censored topics with safety scores intact. TechOrange frames it as the Western answer to Chinese open-source dominance — and cross-checked against the official post and ten outlets, the report is faithful. Our audit adds refinements: GPT-5.5 (not "GPT-5"), the ~1/13 cost holds only vs Claude Opus, the Bridgewater case is a vendor co-publication — and Inkling itself partly learned from Kimi K2.5.

By the Numbers all figures verified vs official post

  • 975B/41Bmultimodal MoE — total / active parameters
  • 1Mmax context on open weights (served tiers 256K / 64K)
  • 0.2–0.99controllable thinking-effort dial
  • 98.6%StrongREJECT — safety intact despite censorship non-compliance

What Follows 起 · 承 · 轉 · 合

BeatPageWhat it covers
起 · 承 The model2Specs, Apache 2.0, the effort dial, CoT condensation, the censorship stance.
轉 The bet & the audit3Adoption economics via Tinker; the Bridgewater case with caveats; figure audit.
合 The frame & its irony4The West-vs-China framing — and the Kimi K2.5 distillation TC omits.

Bottom line

TechOrange's synthesis is faithful — no fabricated figures. Every defect found is a refinement (naming, scope, sourcing), not a debunking.

Reframe

The sharpest omitted fact: the "Western answer" to Chinese open models was itself partly distilled from Kimi K2.5.

Source · TechOrange (廖紹伶, 16 Jul 2026) + TM official post & model card, VentureBeat, TechCrunch, Axios, Fortune, MarkTechPost, Simon Willison, latent.space, the-decoder, Hugging Face. Full per-claim sourcing in the underlying research report.

起 — What Inkling Is released 15 Jul 2026 (US time)

  • Licence · availability

    Open Weights, Apache 2.0

    free commercial use

    • Weights on Hugging Face (thinkingmachines/Inkling) — download, run, customise; no royalties.
    • Runs on SGLang, vLLM, TokenSpeed, llama.cpp; NVFP4 build for Blackwell (≥600 GB).
  • Architecture · scale

    975B Multimodal MoE

    41B active

    • Natively multimodal — text, image, audio reasoning; 975B total / 41B active.
    • 1M-token context on the open weights; Tinker/API tiers serve 256K + 64K.

承 · 轉 — Three Highlights, and the Business Bet dial · condensation · adoption economics

A thinking-effort dial sweeps 0.2–0.99, trading depth for latency. Past 30M RL rollouts the chain of thought spontaneously condensed — "more concise over time, dropping grammatical overhead while remaining comprehensible" — cutting latency sharply. Cognition's Censorship Eval shows "strong patterns of censorship non-compliance" — direct answers where others refuse — while StrongREJECT 98.6% and FORTRESS 78.0% / 95.9% hold. TM concedes open weights allow refusal removal, advising defence-in-depth plus Llama Guard.

  • 0.2–0.99thinking-effort dial — depth vs latency
  • 30M+RL rollouts → spontaneous CoT condensation

TC's core read: Inkling bets on enterprise-adoption economics, not leaderboard rank — monetised through Tinker fine-tuning and managed services, shifting spend from renting closed APIs to owned, customised infrastructure. The showcase: Bridgewater fine-tuned Qwen3-235B on Tinker for financial-document classification — 84.7% vs Claude Opus 78.2% (~$100/1k tasks), GPT-5.5 78.0% (~$70), at ~$7.25/1k. But the result is a joint Bridgewater/TM co-publication — vendor-interested, no independent replication — and one nuance TC softens: Tinker itself is token-metered; the real contrast is owned-custom vs rented-closed-API, not metered vs unmetered.

  • 84.7%fine-tuned Qwen3-235B — vs Opus 78.2 / GPT-5.5 78.0
  • 13.8×cost edge — vs Claude Opus only; ~9.7× vs GPT-5.5

Figure Audit refinements, not debunking

ItemTechOrangeAudit finding
Model name"GPT-5" in the Bridgewater comparisonCorrect name: GPT-5.5 (78.0%, ~$70/1k tasks)
~1/13 costpaired with both GPT-5(.5) and Claude Opusholds vs Opus only (13.8×); ~9.7× vs GPT-5.5
Inkling-Small「2,7?0 億」 — digits cut; active count omitted276B total / 12B active, weights pending

合 — The West-vs-China Frame & Its Irony and what it leaves out

TC asks whether open AI is still "only a Chinese story": with Llama 4 underperforming and Meta closing up, enterprises moved to Qwen, DeepSeek, GLM and Kimi — Inkling is the Western counter-move, with TM's manifesto casting the closed frontier paradigm as "central planning" and citing Hayek on dispersed knowledge. All confirmed. What TC omits is the twist: Inkling's early post-training data was partly distilled from Kimi K2.5 — the Western answer partly learned from the Chinese models it answers. Also unmentioned: day-0 Baseten/Databricks support, and that future TM models are case-by-case — not necessarily open. Where it lands: open-vs-closed is undecided, but the West has pushed a new chip onto the table (「西方陣營至少已經把一張新的籌碼推上了牌桌」) — the chip is real; read its fine print.

  • K2.5Kimi K2.5 — the distillation source TC omits
  • 45Ttokens of multimodal pretraining (also omitted)