Tratopedia
繁中
Settings

Text size

Language

Theme

High contrast

Version

v1.177.0

The release this page was built from. It is what the service worker caches under.

Open Models · Language Models · AI Capability · Geopolitics · Model BriefReleased 15 Jul 2026 · record to 26 Aug 2026

Executive summary · one-page brief

An Open Model, Partly Taught by the Rivals It Answers

Thinking Machines Lab — founded by ex-OpenAI CTO Mira Murati — released its first open-weight model, Inkling, on 15 Jul 2026: a 975B / 41B-active multimodal MoE with a 1M-token open-weight context, Apache 2.0, and a thinking-effort dial. It arrived into a market Chinese labs already dominate by volume — independent measurement puts Chinese open-weight models at roughly 61% of tokens routed on OpenRouter by May 2026, with Meta's own Llama below 1%. Six weeks on, Inkling is independently scored as the leading open-weights release from a U.S. lab. And Thinking Machines' own account of how it was built states plainly that its post-training drew in part on synthetic data from Kimi K2.5 — one of the very models it is measured against.

By the Numbers checked directly against Thinking Machines' own repositories and site

  • 975B/41Bmultimodal MoE — total / active parameters
  • 1Mopen-weight context (Tinker serves 256K / 64K)
  • 41Artificial Analysis Intelligence Index — leading U.S. open-weights score
  • 98.6%StrongREJECT (FORTRESS 78.0% adversarial / 95.9% benign)

What Happened, Graded confirmed at Thinking Machines' own material unless noted otherwise

StandingWhat happenedWhen
ConfirmedInkling released: 975B total / 41B active MoE, 1M-token open-weight context, Apache 2.015 Jul 2026
ConfirmedPost-training bootstrap SFT ran on synthetic data from open-weights models “including Kimi K2.5”training, per TM's own account
ConfirmedChain of thought spontaneously condensed past 30M RL rollouts, unprompted by the rewardduring RL training
Confirmed but not verifiable hereInkling “exhibited strong patterns of censorship non-compliance” on Cognition's Censorship EvalTM's own claim; no independent Cognition score for Inkling found
ConfirmedInkling-Small released with full weights: 276B total / 12B active30 Jul 2026
Confirmed, independentInkling debuts at 41 on Artificial Analysis's Intelligence Index — the leading open-weights score from a U.S. lab15 Jul 2026

Timeline Jun–Aug 2026, each event dated at its own primary source

  1. 24 Jun 2026Independent analysis (Data Gravity) reports Chinese open-weight models at ~61% of OpenRouter's routed tokens by May 2026; Meta's Llama below 1%.
  2. 30 Jun 2026Bridgewater AIA Labs and Thinking Machines publish a Tinker fine-tuning case study — on Qwen3-235B, two weeks before Inkling exists.
  3. 10 Jul 2026Thinking Machines publishes “The Future Worth Building Is Human,” quoting Hayek on dispersed knowledge and warning against concentrating AI development in a few labs — it names no country.
  4. 15 Jul 2026Inkling released; day-0 API access via Together AI, Fireworks, Modal, Databricks and Baseten. Artificial Analysis scores it 41 on its Intelligence Index the same day.
  5. 30 Jul 2026Inkling-Small released with full weights (276B total / 12B active); Artificial Analysis scores it 40, within a point of Inkling at under a third of the parameters.
  6. 31 Jul 2026Thinking Machines publishes its safety-process account: internal testing, four named external red-teams, and an adversarial fine-tuning check on both models.
  7. 26 Aug 2026Hugging Face records 166,108 downloads and 1,749 likes for Inkling, and 168,647 downloads with 379 likes for Inkling-Small.

The chronology matters because two of these events are easy to conflate. The Bridgewater/Tinker case study — a fine-tuned Qwen3-235B beating frontier models on financial-document classification — was published two weeks before Inkling existed. It demonstrates what fine-tuning on Thinking Machines' Tinker platform can do; it says nothing about Inkling's own performance, and no dedicated Inkling deployment case is publicly documented. It is included here because it is the platform-economics evidence behind Inkling's own fine-tuning story, not because it is about Inkling.

What Thinking Machines Says Inkling Is For attributed to Thinking Machines' own material

Thinking Machines states its own case plainly: Inkling is “not the strongest overall model available today, open or closed.” Its bet is a combination — multimodal capability, an efficient controllable thinking-effort dial, and availability on Tinker for fine-tuning — that makes it “a good open-weights base for customization,” not a leaderboard entry. The architecture: a 66-layer decoder-only transformer routing each token to 6 of 256 experts plus 2 always-active shared experts, hybrid local/global attention, and an encoder-free multimodal design — images patch-encoded, audio represented as dMel spectrograms, both projected into the same hidden space the decoder processes text in. The thinking-effort dial's technical range is 0 to 0.99; Thinking Machines' own published benchmark curves sweep it from 0.2 to 0.99, where Inkling reaches a given Terminal Bench 2.1 score at roughly a third the tokens of Nemotron 3 Ultra.

  • 0–0.99thinking-effort dial — full technical range
  • 30M+RL rollouts → spontaneous CoT condensation

The Market It Entered, and the Complication attributed by name, not adopted as Tratopedia's own reading

Two independent measurements frame where Inkling landed. Data Gravity's own token-routing analysis: by May 2026, Chinese open-weight models held ~61% of tokens routed on OpenRouter, “four of the five most-used models are Chinese — and Meta's Llama, the open-weight leader two years ago, has fallen off the rankings entirely,” below 1% of routed volume. Artificial Analysis, the same day Inkling shipped: it is “the leading open weights release from a U.S. lab,” three points above the previous leader, Nemotron 3 Ultra. The complication is that Thinking Machines' own stated rationale does not frame it this way. Its manifesto, published five days before release, argues against concentrating AI development in a “central planning”-like handful of labs, quoting Hayek on dispersed knowledge — and never names China. The “Western answer” framing belongs to the independent market measurements above, not to Thinking Machines' own words. And Inkling's own post-training complicates whichever framing is used: its bootstrap fine-tuning drew on synthetic data from open-weights models “including Kimi K2.5” — one of the very Chinese models the market-share figures name.

The Tinker Case, Re-derived From Its Own Chart not Inkling — Qwen3-235B, fine-tuned two weeks before Inkling shipped

ModelAccuracyCost / 1,000 tasks
Bridgewater's trained model (Qwen3-235B, fine-tuned)84.7%$4.72
GPT 5.578.2%$65.06
Claude Opus 4.878.0%$92.59
Claude Opus 4.677.2%$61.52
GPT 5.475.8%$28.50
Gemini 3.1 Pro74.3%$31.77

The comparison Bridgewater and Thinking Machines publish jointly sets their fine-tuned model against six frontier models. The strongest of them on this task is GPT 5.5, at 78.2% average accuracy and $65.06 per thousand tasks; Claude Opus 4.8 scores 78.0% but costs $92.59, the most expensive in the set. The fine-tuned Qwen3-235B reaches 84.66% at $4.72 — 13.8× cheaper than GPT 5.5, the figure the post itself states, and 19.6× cheaper than Claude Opus 4.8. Two caveats travel with it: the study is a joint publication by the platform's vendor and its customer, with no independent replication, and it measures a Qwen3-235B fine-tune rather than Inkling.

Safety Testing from “A Safe Path to Open Weights,” 31 Jul 2026

  • Internal · three tracks

    Dangerous Knowledge, Misuse, Multimodal Harm

    • CBRN and offensive-cyber knowledge and operational uplift, tested internally.
    • A multimodal harm test across 17 languages and text, image and audio inputs.
  • External · four named testers

    Scale AI, Handshake AI, FAR.AI, Apollo Research

    • Each mandated to a distinct area: general misuse, vulnerable-user harm, CBRN/cyber elicitation, loss-of-control behaviour.
    • None found capability meaningfully beyond what existing open-weight models already pose.

Stripping the Guardrails, Deliberately the adversarial fine-tuning test

Because open weights can be re-tuned to strip refusal training, Thinking Machines fine-tuned “helpful-only” variants of both Inkling and Inkling-Small — deliberately optimised to comply rather than refuse — and ran them against the same dual-use evaluations. Its own conclusion: “the helpful-only variants did not provide new uplift on CBRN and cyber tasks, and remained comparable to existing open-weight models.” Combined with the internal and external results, Thinking Machines concluded releasing Inkling did not add material risk beyond what already-downloadable open-weight models pose.

Inkling-Small, Six Weeks On independent benchmarking, Artificial Analysis

Artificial Analysis's own breakdown, published the day Inkling-Small shipped: it edges ahead of Inkling on Humanity's Last Exam (32% vs 30%) and GPQA Diamond (89% vs 87%) at under a third of the parameters, but trails on agentic work — τ³-Banking 15% vs Inkling's 24% — and scores negative on Artificial Analysis's own factual-calibration index, where Inkling scores positive, “meaning incorrect answers outweigh correct ones.” Both models were independently placed close together on the Intelligence Index (40 vs 41), which is a coarser summary than either of these individual results and does not on its own say which model suits which task.

Conclusion what is settled, and what is not

What Inkling is

A genuine, well-documented open-weight release — 975B/41B MoE, 1M-token open-weight context, Apache 2.0 — independently scored as the leading U.S. open-weights entry six weeks in, on a market Chinese labs already dominate by volume.

What its own account adds

By Thinking Machines' own account, Inkling's post-training was bootstrapped with an initial round of supervised fine-tuning on synthetic data from open-weight models “including Kimi K2.5” — one of the Chinese models it is measured against. The company puts that bootstrap at “a small fraction of compute”, with the bulk going to large-scale reinforcement learning, so this is a starting point rather than the substance of the model.

  • The Bridgewater/Tinker economics case is real and well-documented, but it is a Qwen3-235B fine-tune from before Inkling existed. No dedicated Inkling enterprise deployment is publicly documented — hold that gap, rather than reading the Tinker case as evidence of Inkling's own results.
  • The Cognition Censorship Eval result is Thinking Machines' own claim about its own model. No independently published Cognition score naming Inkling was found — a fact worth knowing before treating the finding as independently verified.
  • “Western answer to Chinese open models” is a reading of independently measured market share (Data Gravity, Artificial Analysis), not Thinking Machines' own stated argument, which never names China. Keep the two apart.

Sources, all fetched directly and captured 2026-08-26 unless a different capture or publish date is stated: Thinking Machines Lab's own Inkling and Inkling-Small release posts, product page, model repositories on Hugging Face (model card, config.json, and live API records), “The Future Worth Building Is Human,” and “A Safe Path to Open Weights” · the joint Bridgewater AIA Labs/Thinking Machines post on Tinker fine-tuning economics, read for its own embedded chart data · Artificial Analysis's two independent launch-day articles · “China's Open-Weight Takeover” (Chris Zeoli, Data Gravity) · Cognition's own methodology post, read to check for an Inkling-specific score.

Versions

This document is rewritten when what it says has to change. Every version stays published at its own address.

  1. v0001 current

The current version is also at latest/.