Tratopedia
繁中
Settings

Text size

Language

Theme

High contrast

Version

v1.177.0

The release this page was built from. It is what the service worker caches under.

Local inference · Language Models · Apple Silicon · AI Agents · Technical reportGemma 4 of April 2026, record to August 2026

Gemma 4 · Apple Silicon · Claude Code and Copilot CLI, four months on

The real bottleneck is KV cache and prefill, not the context limit

This report works out what it takes to run Google's Gemma 4 locally on a Mac, behind an agentic CLI tool such as Claude Code or GitHub Copilot CLI. The conclusion first: what sets how an agent feels to use is prefill speed and KV-cache footprint, not how much context a model can accept — which is also why Gemma 4's 26B-A4B mixture-of-experts model, with only 3.8B parameters active per token, is the sweet spot for local agentic work. Since April 2026 Apple has raised Mac prices twice, and two Ollama bugs now bear on which Gemma 4 size is actually safe to run.

  • 3.8Bactive parameters in the 26B-A4B model recommended below, out of 25.2B total
  • 256Kmaximum context on Gemma 4's three larger models — 12B, 26B-A4B and 31B
  • +35%rise in Taiwan's entry Mac mini price since April 2026
  • 614 GB/sunified-memory bandwidth on Apple's M5 Max, at 128GB

About the prices

Every price below is Apple's own Taiwan retail price as listed on 25 August 2026, or a US price as reported on the date given beside it. This is a year in which memory and storage costs have moved sharply — Apple raised Mac mini prices twice between May and June alone — so check the current figure before buying rather than trusting one printed here. None of this is purchasing advice.

What's changed since April Graded by how it's confirmed, strongest first

StandingWhat happenedDetail
ConfirmedGemma 4 replaces Gemma 3's namingGoogle DeepMind released Gemma 4 on 2 April 2026 in five sizes — E2B, E4B, 12B, 26B-A4B and 31B — replacing Gemma 3's 1B/4B/12B/27B lineup from March 2025.
ConfirmedApple raised Mac mini prices twiceThe entry configuration's effective US price rose from $599 to $799 in May 2026, and the M4 Pro's starting price rose from $1,399 to $1,599 in June. Apple told the Wall Street Journal, quoted by MacRumors, that it was the sharpest component-cost rise it had seen, tied to AI-datacentre demand for memory and storage.
ConfirmedBoth major coding CLIs now reach local modelsOllama added a native, Anthropic-compatible API in January 2026, so Claude Code can point at a local server directly. GitHub Copilot CLI added bring-your-own-key support, including local models, in June 2026.
Confirmed, narrower than first reportedA flash-attention hang hits one Gemma 4 size, not the one recommended hereOllama users on both Nvidia CUDA and Apple Silicon report Gemma 4's 31B Dense model hanging indefinitely under Flash Attention on long prompts. At least one report states the 26B-A4B model this page recommends is not affected by the same prompts.
Confirmed, still openOllama's Qwen 3.5 27B tool calling remains brokenA repetition-penalty bug was patched in v0.17.5, but the deeper fault — Ollama's correct tool-calling renderer wired to the wrong model name — was still open as of the report found.
Reported, not independently verifiedM5 Max is faster for local inference, but by how much is unclearApple's own announcement states GPU-compute and graphics multiples, not tokens per second; the only LLM throughput figures found are published by a benchmark blog that labels its own numbers “Estimated.”

Four Years in Chips and Models Background first, then the year this report covers

  1. 2023Apple cuts the M3 Pro's memory bandwidth to 150GB/s, from 200GB/s on the M2 Pro — background worth knowing, since Mac buyers often compare across chip generations.
  2. 2024.10Mac mini M4 and M4 Pro launch; Taiwan's entry configuration (16GB/256GB) costs NT$19,900.
  3. 2025.03Google DeepMind releases Gemma 3, in 1B/4B/12B/27B sizes.
  4. 2026.01Ollama ships a native, Anthropic-compatible API (v0.14), letting Claude Code run directly against a local model.
  5. 2026.03Apple announces the M5 Pro and M5 Max, with unified-memory bandwidth up to 614GB/s.
  6. 2026.04.02Google DeepMind releases Gemma 4 — E2B, E4B, 12B, 26B-A4B, 31B.
  7. 2026.04Ollama users on CUDA and, separately, on Apple Silicon report Gemma 4's 31B Dense model hanging under Flash Attention on long prompts.
  8. 2026.05.01Apple discontinues the Mac mini's cheapest configuration; the entry price effectively rises from $599 to $799.
  9. 2026.06.17GitHub Copilot CLI adds bring-your-own-key support for local and external models.
  10. 2026.06.25Apple raises the Mac mini M4 Pro's starting price from $1,399 to $1,599, and reinstates the discontinued entry configuration at the new, higher price.
  11. 2026.08.25Apple's Taiwan store lists the Mac mini from NT$26,900, and its M4 Pro configurations at NT$54,900 and NT$61,900.

Why context length is the wrong number to watch Prefill and KV cache, not the token limit

An agentic CLI tool re-sends its entire working state — a system prompt, a tool schema, and the running transcript — on every single turn. For a coding agent that can be several thousand tokens before the model produces a single new word, and re-processing that state, called prefill, is what a Mac's memory bandwidth actually has to keep up with. Context length only bounds how long a session can run; it says nothing about how that session feels, turn by turn. What sets that feel is prefill speed, and how much memory the KV cache — the running memory of everything already processed — occupies alongside the model's own weights.

  • 1,024tokens: the sliding-window size bounding most of Gemma 4's attention layers, per the model's own published specification.
  • ×0.5Quantising the KV cache from 16-bit floating point to 8-bit integers halves its memory footprint outright — a definitional fact about those formats, not a measured benchmark.

The naming everyone gets wrong Gemma 4 is not 1B/4B/12B/27B

Commonly saidGemma 4 SKUActive parametersMax context
1BE2B~2.3B (reported)128K
4BE4B~4.5B (reported)128K
12BNo equivalent; nearest is 26B-A4B3.8B of 25.2B total (confirmed)256K
27B31B dense31B (dense, all active)256K

The 26B-A4B row is confirmed directly from the model's own published card: 25.2B total parameters, 3.8B active per token through eight of a hundred and twenty-eight experts (plus one always-on shared expert). The E2B/E4B “effective parameter” figures are reported by aggregated model listings rather than confirmed on a single page this research opened directly, and are marked accordingly. None of these four rows is a size-for-size replacement for its Gemma 3 counterpart — the table exists to correct the mapping, not to declare one Gemma 4 model equivalent to one Gemma 3 model.

Known frictions on Ollama and Apple Silicon Confirmed, but narrower than first reported

  • 31B Dense only

    The flash-attention hang

    • Reported on Nvidia CUDA hardware (an RTX 3090) for prompts over roughly 3,000–4,000 tokens, and separately on Apple Silicon (an M5 Max) for prompts over roughly 500 tokens.
    • The reported root cause: Gemma 4 mixes roughly fifty sliding-window attention layers with ten global-attention layers, and the two use different head dimensions (256 versus 512) that Ollama's Flash Attention implementation did not originally handle.
    • At least one report states plainly that the 26B-A4B model — the one this page recommends — handles the same long prompts without issue; the bug is specific to 31B Dense.
  • Still open

    Qwen 3.5 27B tool calling

    • Three separate faults were reported together: sampling penalties silently ignored, an unclosed <think> tag corrupting later turns, and Ollama's correct tool-calling renderer wired only to the model name “qwen3-coder,” not “qwen3.5.”
    • The sampling-penalty bug was patched in v0.17.5; the renderer mismatch was still open as of the report found.
  • Not yet accelerated

    Gemma 4 runs without MLX on Apple Silicon

    • Ollama's Apple Silicon builds normally use Apple's MLX framework for speed; one report found Gemma 4 falling back to the slower llama.cpp backend instead, at roughly 15 tokens per second on 31B Dense and roughly 75 on 26B-A4B.
    • No MLX-accelerated figure for Gemma 4 on Apple Silicon was found in this research; if and when one ships, both numbers should rise.

From the M3 Pro to the M5 Max Three years of the same trade-off

Apple's Silicon line has cut corners on memory bandwidth before: the M3 Pro shipped in October 2023 with 150GB/s, a 25% cut from the M2 Pro's 200GB/s, on a narrower memory bus — a reminder that a newer chip number does not automatically mean more bandwidth for a model that has to stream its own weights past a bottleneck around the clock. The M5 Pro and M5 Max, announced March 2026, move the other way: a new two-die “Fusion Architecture,” more than four times the GPU compute of the previous generation on Apple's own figures, and unified-memory bandwidth up to 614GB/s at the top 128GB configuration. A benchmark blog's own, self-labelled “Estimated” testing puts the M5 Max ahead of the M5 Pro and of a reference RTX 4090 on several model sizes, though Apple's own announcement does not itself state a tokens-per-second figure for any model.

Two CLIs, two ways in Different protocols, different timelines

ItemClaude CodeCopilot CLI
Native local-model supportYes, since Ollama v0.14 (Jan 2026)Yes, bring-your-own-key, since June 2026
ProtocolAnthropic Messages APIOpenAI-compatible / external providers
Set-upPoint ANTHROPIC_BASE_URL at a local Ollama serverConfigure a local or external model in the CLI's model picker

Which Mac, for what Fit, not a winner

  • Entry point

    Learning the ropes on a small model

    • Who it's for: someone new to local inference who wants to try a small agent — an 8B-class dense model — without committing much money.
    • Best at: Apple's own entry Mac mini (16GB/256GB, NT$26,900 as of August 2026) is quiet, sips power, and needs nothing beyond stock Ollama to get going.
    • Trade-off: Apple prices build-to-order memory and storage through its configurator rather than as a listed figure, so budget a premium over the entry tier rather than a fixed number.
  • The sweet spot

    Gemma 4's 26B-A4B, comfortably

    • Who it's for: someone who wants the specific model this report argues is the local-agentic sweet spot, with headroom for its 256K context.
    • Best at: more unified memory (24GB or more) is what actually buys agent responsiveness here, because it is what lets the model's active parameters and its KV cache sit alongside the OS without swapping.
    • Trade-off: Apple prices build-to-order memory and storage through its configurator rather than as a listed figure, so budget a premium over the entry tier rather than a fixed number.
  • Maximum local capability

    M5 Max, if the budget stretches

    • Who it's for: someone who wants to run the largest models locally and is prepared to pay for it.
    • Best at: up to 614GB/s of unified-memory bandwidth and more than four times the previous generation's GPU compute, on Apple's own published figures.
    • Trade-off: it currently ships only in a MacBook Pro chassis, where sustained inference risks thermal throttling; a benchmark blog reports a desktop Mac Studio equivalent as not yet available.
  • Not on a Mac at all

    An NVIDIA machine, or the cloud

    • Who it's for: someone who already owns capable NVIDIA hardware, or who doesn't want to manage Ollama versions and KV-cache flags at all.
    • Best at: an NVIDIA card's VRAM is dedicated to the model rather than shared with the OS, and a cloud API removes version and configuration management entirely.
    • Trade-off: continuous NVIDIA-GPU power draw is well known to run far higher than an Apple Silicon Mac's, and neither option is private or free per token the way a Mac already paid for is.

Bottom line

What changed since April is not the argument — KV cache and prefill still set an agent's feel more than context length does — but the price of acting on it. Apple's own Mac mini now costs meaningfully more than it did when this report first ran, the specific mid-tier configuration this report used to name a price for could not be re-confirmed, and two real Ollama bugs now bear on which Gemma 4 size is actually safe to run. None of that changes which model to run; it changes what it costs, and it is worth checking the current price for yourself before buying.

Sources: Google Developers Blog and Android Developers Blog; Hugging Face (google/gemma-4-26B-A4B; Unsloth's Gemma 4 GGUF listings); MacRumors (25 Jun 2026, 1 May 2026, 31 Oct 2023); NewMobileLife (26 Jun 2026); Apple's Taiwan store, fetched 25 Aug 2026; PChome 24h 購物; kocpc.com.tw; GitHub issues #15350, #15368 and #14493 on ollama/ollama, read only via search; GitHub's own changelog (17 Jun 2026); Apple Newsroom (M5 Pro/M5 Max, 3 Mar 2026); PromptQuorum's self-labelled “Estimated” benchmark page. None of this is purchasing advice.

You are reading v0001, published 2026-05-01. It has been superseded — the current version is v0002.

Versions

This document is rewritten when what it says has to change. Every version stays published at its own address.

  1. v0002 current
  2. v0001 superseded

The current version is also at latest/.