Skip to content

AI compute

Where should your inference run?

A Mac you already own generates tokens for the price of the electricity. A GPU node generates them faster than any Mac ever will. A cloud API generates better ones than either. The answer is not which is cheapest. It is which of the three your workload actually disqualifies.

Prices last checked 2026-08-06. Hardware and token prices were verified on 2026-08-06 and move fast. Apple raised prices across the Mac range on 2026-06-25, the Mac Studio M3 Ultra by 32.5%, other models by different amounts, and withdrew two Mac Studio memory tiers earlier in the year; Anthropic raises Sonnet 5 by 50% on 2026-09-01; H100 reserved rental rose ~40% off its October 2025 low.

Your fleet

The hardware is already paid for. One input below is pure judgment, it says so.

One machine per person, each user's Mac does that user's own routable work.

Throughput comes from the published chip-level benchmark; capex is sunk and never charged.

What people do with AI

28,000 output tokens per person per day, derived, not observed.

Share routable on-device25%

Your judgment, not a fact, nobody publishes this. Routable: classification, extraction, summarising a known document, redaction, autocomplete. Not routable: frontier reasoning, long-context analysis, agentic coding.

Routable tasks belong on a cheap tier, displacing a frontier price with them would inflate the saving.

Prompt cache hit rate50%

Repetitive routable work caches well, a higher rate shrinks the API bill being displaced, so it is the conservative direction.


Refresh cycle

Gartner puts enterprise laptops at 3.7 years on average (2024, reported second-hand), the devices get replaced on this clock whether or not AI exists.

The only recurring hardware cost in this story: the memory headroom you choose at a purchase that was happening anyway.

The fleet you already own

Routing 25% of this work on-device displaces $11,594 of API spend a year, on hardware that was bought anyway.

500 × MacBook Air 13" M4 (16GB): typical fleet device · capex sunk, refreshed every 4 years regardless · the 25% routable share is your judgment, not a measurement.

Displaced API spend

$11,594/yr

vs paying Claude Haiku 4.5 for all of it

On-device cost

≈ $0

floor, inference draw unmeasured on this device, so no electricity line is invented

Refresh-cycle delta

$4/device/mo

MacBook Air M5: 16 → 24GB amortised over 4 years

Does each machine keep up?

  • The routable share is 7k tokens per person per day, about 13 minutes of inference at this machine's effective 9.2 tok/s. The ceiling accepted here is 2 hours a day (25% duty on an 8-hour day), because laptops throttle under sustained load.

  • Effective speed is below comfortable chat streaming (15 tok/s) once prompt-reading is paid for, fine for background work (classification, redaction, summarising), wrong for a chat window someone watches. The routable list is background work on purpose.

  • What this is not: a serving cluster. Apple runtimes do not batch concurrent users and a fanless laptop sheds ~55% of its package power after 20 minutes of sustained load (Notebookcheck, M4 Air). Each Mac carries its own user, nothing more.

The annual bill, three ways

  • All tokens on Claude Haiku 4.5$46,375/yr

    The do-nothing baseline.

  • Hybrid: 25% on-device, rest on the API$34,781/yr

    On-device leg priced at an unmeasured $0 floor, the true number is above this by whatever the electricity turns out to be.

  • Per-seat assistant for everyone (M365 Copilot)$180,000/yr

    US$30/user/month list, a different product with a different ceiling, shown for scale.

The refresh question

Speccing MacBook Air M5: 16 → 24GB at the next refresh costs $4/device/month over a 4-year cycle, against US$30/user/month for a Copilot seat ($26/month less). They are not substitutes: the RAM buys headroom for on-device routable work, the seat buys a frontier assistant. The point is scale, the entire hardware cost of this story is smaller than one seat licence.

What this model will not claim

Each of these is a conclusion the evidence did not support, and each one is routinely asserted elsewhere.

  • Macs replace cloud LLM APIs at scale.

    They do not. No published, named-enterprise case study of a Mac fleet replacing cloud LLM APIs survived verification. The one vendor case study found specifies a "Mac Studio M4 Ultra", a chip Apple has never shipped.

  • AI governance and local inference are the same argument.

    Jamf AI Governance controls cloud AI tools running on Macs. It neither requires nor evidences local inference. Conflating them is the fastest way to lose a procurement review.

  • Regulation mandates on-device inference.

    No regulator we found requires it. The closest real forcing function is that HHS OCR treats LLM providers handling PHI as business associates, so a BAA is needed, local inference sidesteps that contracting exercise.

  • Apple's on-device model can replace an enterprise assistant seat.

    The on-device Apple foundation model is roughly 3B parameters with an 8,192-token context window. Private Cloud Compute extends context to 32k. Neither is a Copilot substitute.

  • Buying an 8x H100 node is a sensible 2026 purchase.

    Owned H100 at 100% duty produces tokens at roughly $0.165/M. Rented Blackwell produces them at roughly $0.081/M. H100 is two generations behind the shipping frontier.

Numbers we could not source

  • Tokens consumed per knowledge worker per day

    No credible published data exists. Vendors publish dollars, seats and growth rates. The defaults here are derived from Anthropic's published per-user spend divided by a stated model price, and are labelled as derived wherever they appear.

  • Sysadmin cost per machine per year

    No sourced figure exists for a Mac fleet or a GPU node. The input starts at $0 so the omission is visible rather than silent.

  • The dollar value of privacy, residency or compliance risk

    None is defensible, so privacy removes options rather than pricing them. The single exception is Anthropic's US-residency surcharge of 1.1x on every token category, the only published price for a residency guarantee we found.

  • Open-weight model API prices (Together, Fireworks, Groq, DeepInfra)

    Not researched, and this is the most important missing comparator: hosted open-weight inference is the correct baseline for "should we self-host". Comparing against a frontier API instead favours the local option.

  • Apple hardware residual value

    Not researched. Local hardware is amortised to zero, which understates the Mac case rather than flattering it.

  • Concurrency benchmarks for a 70B+ model on any Mac

    None exists at any batch size above one. Aggregate throughput is assumed flat on the strength of a 9B measurement across four backends, where MLX, llama.cpp and Ollama were all unchanged from one to four concurrent requests. A 0.6B model on an M4 Max did scale 3.7x from 1 to 16, so the flat assumption is drawn from the larger model, and a small enough one may behave differently.

How to read a Mac benchmark

Identical hardware running an identical model spans 3× on token rate depending only on the runtime, MLX 130, llama.cpp 89, Ollama 44 tokens/sec on one measured comparison. A figure without a named tester, runtime version, quantisation, context length and batch size cannot be checked and should be assumed fabricated. A cluster of sites currently publishes precise-looking tables for a “Mac Studio M5 Max” and an “M4 Ultra”, neither of which Apple has ever shipped. Every benchmark on this page names its tester, runtime, quantisation and batch size, and links its source. Where a source did not disclose the context length, the row says so rather than inventing one, and where a source could not be retrieved at all, the benchmark was removed rather than cited.

Compiled 2026-08-06 from vendor pricing pages, published independent benchmarks, EIA electricity statistics, CBRE colocation data and the Uptime Institute data centre survey. Every figure carries a source, a stated derivation, or an admission that it is missing. Two conventions are neither: a 250-day working year and a three-year amortisation for local hardware. Both scale every annual figure on this page and both are choices, not findings.

Where every number comes from

All 53 figures this page can use, each with a link and a label saying what kind of evidence it is. A manufacturer's rating and somebody's actual measurement are not the same thing, so they are never shown as though they were. Last checked 2026-08-06.

List price
The published list price. Most buyers negotiate below it.
Measured
A third party measured it and published their method.
Measured once
One party measured it. Nobody has replicated it yet.
Reseller price
A reseller's price, because the vendor does not publish one.
Our assumption
A convention we chose, not a finding. Change it if yours differs.
Official statistic
A government or official statistical series.
Rental rate
A published hourly rental rate.
Our arithmetic
Our own arithmetic on sourced inputs. The working is shown.

Validate your numbers with a fleet expert

30 minutes with an Apple fleet consultant: pressure-test the assumptions, map your rollout risks, leave with next steps. Free, no pitch.