Tratopedia
繁中
Settings

Text size

Language

Theme

High contrast

Version

v1.177.0

The release this page was built from. It is what the service worker caches under.

Local inference · Language Models · Apple Silicon · AI Agents · Technical reportGemma 4 of April 2026, record to September 2026

Gemma 4 · Apple Silicon · Claude Code and Copilot CLI, four months on

The real bottleneck is KV cache and prefill, not the context limit

This report works out what it takes to run Google's Gemma 4 locally on a Mac, behind an agentic CLI tool such as Claude Code or GitHub Copilot CLI. The conclusion first: what sets how an agent feels to use is prefill speed and KV-cache footprint, not how much context a model can accept — which is also why Gemma 4's 26B-A4B mixture-of-experts model, with only 3.8B parameters active per token, is the sweet spot for local agentic work. Since April 2026 Apple has raised Mac prices twice, and two Ollama bugs now bear on which Gemma 4 size is actually safe to run.

  • 3.8Bactive parameters in the 26B-A4B model recommended below, out of 25.2B total
  • 256Kmaximum context on Gemma 4's three larger models — 12B, 26B-A4B and 31B
  • +35%rise in Taiwan's entry Mac mini price since April 2026
  • 614 GB/sunified-memory bandwidth on Apple's M5 Max, at 128GB

About the prices

Every price below is Apple's own Taiwan retail price as listed on 25 August 2026, or a US price as reported on the date given beside it. This is a year in which memory and storage costs have moved sharply — Apple raised Mac mini prices twice between May and June alone — so check the current figure before buying rather than trusting one printed here. Apple put a new Mac mini, with M6 or M5 Pro, on sale on 22 September 2026; the Mac mini prices and recommendations below are those of the M4 line in August 2026. None of this is purchasing advice.

What 2026 has changed Graded by how it's confirmed, strongest first

StandingWhat happenedDetail
ConfirmedGemma 4 replaces Gemma 3's namingGoogle DeepMind released Gemma 4 on 2 April 2026 in four sizes — E2B, E4B, 26B-A4B and 31B — and added a fifth, the 12B, on 3 June 2026; it replaces Gemma 3's 1B/4B/12B/27B lineup from March 2025.
ConfirmedApple raised Mac mini prices twiceThe entry configuration's effective US price rose from $599 to $799 in May 2026, and the M4 Pro's starting price rose from $1,399 to $1,599 in June. Apple told the Wall Street Journal, quoted by MacRumors, that it was the sharpest component-cost rise it had seen, tied to AI-datacentre demand for memory and storage.
ConfirmedBoth major coding CLIs now reach local modelsOllama added a native, Anthropic-compatible API in January 2026, so Claude Code can point at a local server directly. GitHub Copilot CLI added bring-your-own-key support, including local models, in June 2026.
Confirmed, one model size onlyA flash-attention hang hits one Gemma 4 size, not the one recommended hereOllama users on both Nvidia CUDA and Apple Silicon report Gemma 4's 31B Dense model hanging indefinitely under Flash Attention on long prompts. At least one report states the 26B-A4B model this page recommends is not affected by the same prompts.
Confirmed, still openOllama's Qwen 3.5 27B tool calling remains brokenA repetition-penalty bug was patched in v0.17.5, but the deeper fault — Ollama's correct tool-calling renderer wired to the wrong model name — was still open when last reported.
Reported, not independently verifiedM5 Max is faster for local inference, but by how much is unclearApple's own announcement states GPU-compute and graphics multiples, not tokens per second; LLM throughput figures come instead from a benchmark blog that labels its own numbers “Estimated.”

Four Years in Chips and Models Background first, then the year this report covers

  1. 2023Apple cuts the M3 Pro's memory bandwidth to 150GB/s, from 200GB/s on the M2 Pro — background worth knowing, since Mac buyers often compare across chip generations.
  2. 2024.10Mac mini M4 and M4 Pro launch; Taiwan's entry configuration (16GB/256GB) costs NT$19,900.
  3. 2025.03Google DeepMind releases Gemma 3, in 1B/4B/12B/27B sizes.
  4. 2026.01Ollama ships a native, Anthropic-compatible API (v0.14), letting Claude Code run directly against a local model.
  5. 2026.03Apple announces the M5 Pro and M5 Max, with unified-memory bandwidth up to 614GB/s.
  6. 2026.04.02Google DeepMind releases Gemma 4 in four sizes — E2B, E4B, 26B-A4B, 31B.
  7. 2026.04Ollama users on CUDA and, separately, on Apple Silicon report Gemma 4's 31B Dense model hanging under Flash Attention on long prompts.
  8. 2026.05.01Apple discontinues the Mac mini's cheapest configuration; the entry price effectively rises from $599 to $799.
  9. 2026.06.03Google adds a fifth Gemma 4 size, the 12B “Unified”, which has no separate image or audio encoder.
  10. 2026.06.17GitHub Copilot CLI adds bring-your-own-key support for local and external models.
  11. 2026.06.25Apple raises the Mac mini M4 Pro's starting price from $1,399 to $1,599, and reinstates the discontinued entry configuration at the new, higher price.
  12. 2026.08.25Apple's Taiwan store lists the Mac mini from NT$26,900, and its M4 Pro configurations at NT$54,900 and NT$61,900.
  13. 2026.09.22Apple puts a new Mac mini, with M6 or M5 Pro, and a Mac Studio with M5 Max or M5 Ultra on sale.

Why context length is the wrong number to watch Prefill and KV cache, not the token limit

An agentic CLI tool re-sends its entire working state — a system prompt, a tool schema, and the running transcript — on every single turn. For a coding agent that can be several thousand tokens before the model produces a single new word, and re-processing that state, called prefill, is what a Mac's memory bandwidth actually has to keep up with. Context length only bounds how long a session can run; it says nothing about how that session feels, turn by turn. What sets that feel is prefill speed, and how much memory the KV cache — the running memory of everything already processed — occupies alongside the model's own weights.

  • 1,024tokens: the sliding-window size bounding most of Gemma 4's attention layers, per the model's own published specification.
  • ×0.5Quantising the KV cache from 16-bit floating point to 8-bit integers halves its memory footprint outright — a definitional fact about those formats, not a measured benchmark.

The naming everyone gets wrong Gemma 4 is not 1B/4B/12B/27B

Commonly saidGemma 4 SKUParametersMax context
1BE2B2.3B effective (5.1B with embeddings)128K
4BE4B4.5B effective (8B with embeddings)128K
12B12B Unified (dense)11.95B (dense, all active)256K
No Gemma 3 counterpart26B-A4B (mixture of experts)3.8B active of 25.2B total256K
27B31B dense30.7B (dense, all active)256K

Every row is read off Google's own model cards. In E2B and E4B the “E” stands for effective parameters: each layer carries its own per-layer embedding table, which is large but used only for quick look-ups, so the effective count (2.3B, 4.5B) is much smaller than the total with embeddings (5.1B, 8B). In 26B-A4B the “A” stands for active: 3.8B of its 25.2B parameters work on each token, through eight of a hundred and twenty-eight experts plus one always-on shared expert. The 12B keeps a Gemma 3 size but not its design — “Unified” means it has no separate image or audio encoder, and feeds both straight into the language model. None of these rows is a like-for-like replacement for its Gemma 3 counterpart — the table exists to correct the mapping, not to declare one Gemma 4 model equivalent to one Gemma 3 model.

Known frictions on Ollama and Apple Silicon Real, and each narrower than it sounds

  • 31B Dense only

    The flash-attention hang

    • Reported on Nvidia CUDA hardware (an RTX 3090) for prompts over roughly 3,000–4,000 tokens, and separately on Apple Silicon (an M5 Max) for prompts over roughly 500 tokens.
    • The reported root cause: Gemma 4 mixes roughly fifty sliding-window attention layers with ten global-attention layers, and the two use different head dimensions (256 versus 512) that Ollama's Flash Attention implementation did not originally handle.
    • At least one report states plainly that the 26B-A4B model — the one this page recommends — handles the same long prompts without issue; the bug is specific to 31B Dense.
  • Still open

    Qwen 3.5 27B tool calling

    • Three separate faults were reported together: sampling penalties silently ignored, an unclosed <think> tag corrupting later turns, and Ollama's correct tool-calling renderer wired only to the model name “qwen3-coder,” not “qwen3.5.”
    • The sampling-penalty bug was patched in v0.17.5; the renderer mismatch was still open when last reported.
  • Not yet accelerated

    Gemma 4 runs without MLX on Apple Silicon

    • Ollama's Apple Silicon builds normally use Apple's MLX framework for speed; one report found Gemma 4 falling back to the slower llama.cpp backend instead, at roughly 15 tokens per second on 31B Dense and roughly 75 on 26B-A4B.
    • There is as yet no MLX-accelerated figure for Gemma 4 on Apple Silicon to set beside these; if and when MLX support ships, both numbers should rise.

From the M3 Pro to the M5 Max Three years of the same trade-off

Apple's Silicon line has cut corners on memory bandwidth before: the M3 Pro shipped in October 2023 with 150GB/s, a 25% cut from the M2 Pro's 200GB/s, on a narrower memory bus — a reminder that a newer chip number does not automatically mean more bandwidth for a model that has to stream its own weights past a bottleneck around the clock. The M5 Pro and M5 Max, announced March 2026, move the other way: a new two-die “Fusion Architecture,” more than four times the GPU compute of the previous generation on Apple's own figures, and unified-memory bandwidth up to 614GB/s at the top 128GB configuration. A benchmark blog's own, self-labelled “Estimated” testing puts the M5 Max ahead of the M5 Pro and of a reference RTX 4090 on several model sizes, though Apple's own announcement does not itself state a tokens-per-second figure for any model.

Two CLIs, two ways in Different protocols, different timelines

ItemClaude CodeCopilot CLI
Native local-model supportYes, since Ollama v0.14 (Jan 2026)Yes, bring-your-own-key, since June 2026
ProtocolAnthropic Messages APIOpenAI-compatible / external providers
Set-upPoint ANTHROPIC_BASE_URL at a local Ollama serverConfigure a local or external model in the CLI's model picker

Which Mac, for what Fit, not a winner

  • Entry point

    Learning the ropes on a small model

    • Who it's for: someone new to local inference who wants to try a small agent — an 8B-class dense model — without committing much money.
    • Best at: Apple's M4 entry Mac mini (16GB/256GB, NT$26,900 at Apple Taiwan in August 2026) is quiet, sips power, and needs nothing beyond stock Ollama to get going. That is the August 2026 line: a new Mac mini, with M6 or M5 Pro, went on sale on 22 September 2026.
    • Trade-off: Apple prices build-to-order memory and storage through its configurator rather than as a listed figure, so budget a premium over the entry tier rather than a fixed number.
  • The sweet spot

    Gemma 4's 26B-A4B, comfortably

    • Who it's for: someone who wants the specific model this report argues is the local-agentic sweet spot, with headroom for its 256K context.
    • Best at: more unified memory (24GB or more) is what actually buys agent responsiveness here, because it is what lets the model's active parameters and its KV cache sit alongside the OS without swapping.
    • Trade-off: Apple prices build-to-order memory and storage through its configurator rather than as a listed figure, so budget a premium over the entry tier rather than a fixed number.
  • Maximum local capability

    M5 Max, if the budget stretches

    • Who it's for: someone who wants to run the largest models locally and is prepared to pay for it.
    • Best at: up to 614GB/s of unified-memory bandwidth and more than four times the previous generation's GPU compute, on Apple's own published figures.
    • Trade-off: Apple introduced M5 Max in March 2026 for the MacBook Pro, a laptop; the desktop option, a Mac Studio with M5 Max, went on sale only on 22 September 2026, from US$2,499 in the US.
  • Not on a Mac at all

    An NVIDIA machine, or the cloud

    • Who it's for: someone who already owns capable NVIDIA hardware, or who doesn't want to manage Ollama versions and KV-cache flags at all.
    • Best at: an NVIDIA card's VRAM is dedicated to the model rather than shared with the OS, and a cloud API removes version and configuration management entirely.
    • Trade-off: power. NVIDIA rates the RTX 4090 card alone at 450W of total graphics power and 19W at idle, and asks for an 850W system power supply; Apple measures its base Mac Studio with M5 Max, the whole computer at the wall, at 200W maximum and 7W idle. And a cloud API is neither private nor free per token, the way a machine the buyer already owns is.

Bottom line

What sets an agent's feel is unchanged — KV cache and prefill still matter more than context length — but the price of acting on it has moved. Apple raised Mac mini prices twice in May and June 2026, it prices build-to-order memory upgrades only through its configurator, and two real Ollama bugs bear on which Gemma 4 size is actually safe to run. None of that changes which model to run; it changes what it costs, and it is worth checking the current price for yourself before buying.

What could not be verified

One claim could not be confirmed when searched for on 28 Sep 2026, so this page does not assert it. “Sustained inference risks thermal throttling”, said of the M5 Max in a MacBook Pro chassis — Unconfirmed report: a published troubleshooting guide says every laptop chassis eventually slows under continuous inference, so the idea has an identifiable teller, but no published measurement of throttling on an M5 Max MacBook Pro was found.

Sources: Google's Keyword blog (Gemma 4, 2 Apr 2026; Gemma 4 12B, 3 Jun 2026); Google Developers Blog and Android Developers Blog; Simon Willison's weblog (12 Mar 2025); Hugging Face (google/gemma-4-E2B, -E4B, -12B and -26B-A4B; Unsloth's Gemma 4 GGUF listings); MacRumors (25 Jun 2026, 1 May 2026, 31 Oct 2023); NewMobileLife (26 Jun 2026); Apple's Taiwan store, fetched 25 Aug 2026; PChome 24h 購物; kocpc.com.tw; GitHub issues #15350, #15368 and #14493 on ollama/ollama, read only via search; Ollama's blog (16 Jan 2026) and v0.14.0 release notes; GitHub's own changelog (17 Jun 2026); Apple Newsroom (M5 Pro/M5 Max, 3 Mar 2026; Mac mini and Mac Studio, 22 Sep 2026); Apple Support (Mac Studio power consumption, 21 Sep 2026); NVIDIA (GeForce RTX 4090 specifications); PromptQuorum's self-labelled “Estimated” benchmark page. None of this is purchasing advice.

Versions

This document is rewritten when what it says has to change. Every version stays published at its own address.

  1. v0002 current
  2. v0001 superseded

The current version is also at latest/.