AI compute
Where should your inference run?
A Mac you already own generates tokens for the price of the electricity. A GPU node generates them faster than any Mac ever will. A cloud API generates better ones than either. The answer is not which is cheapest. It is which of the three your workload actually disqualifies.
Prices last checked 2026-08-06. Hardware and token prices were verified on 2026-08-06 and move fast. Apple raised prices across the Mac range on 2026-06-25, the Mac Studio M3 Ultra by 32.5%, other models by different amounts, and withdrew two Mac Studio memory tiers earlier in the year; Anthropic raises Sonnet 5 by 50% on 2026-09-01; H100 reserved rental rose ~40% off its October 2025 low.
Your fleet
The hardware is already paid for. One input below is pure judgment, it says so.
One machine per person, each user's Mac does that user's own routable work.
Throughput comes from the published chip-level benchmark; capex is sunk and never charged.
28,000 output tokens per person per day, derived, not observed.
Your judgment, not a fact, nobody publishes this. Routable: classification, extraction, summarising a known document, redaction, autocomplete. Not routable: frontier reasoning, long-context analysis, agentic coding.
Routable tasks belong on a cheap tier, displacing a frontier price with them would inflate the saving.
Repetitive routable work caches well, a higher rate shrinks the API bill being displaced, so it is the conservative direction.
Gartner puts enterprise laptops at 3.7 years on average (2024, reported second-hand), the devices get replaced on this clock whether or not AI exists.
The only recurring hardware cost in this story: the memory headroom you choose at a purchase that was happening anyway.
The fleet you already own
Routing 25% of this work on-device displaces $11,594 of API spend a year, on hardware that was bought anyway.
500 × MacBook Air 13" M4 (16GB): typical fleet device · capex sunk, refreshed every 4 years regardless · the 25% routable share is your judgment, not a measurement.
Displaced API spend
$11,594/yr
vs paying Claude Haiku 4.5 for all of it
On-device cost
≈ $0
floor, inference draw unmeasured on this device, so no electricity line is invented
Refresh-cycle delta
$4/device/mo
MacBook Air M5: 16 → 24GB amortised over 4 years
Does each machine keep up?
The routable share is 7k tokens per person per day, about 13 minutes of inference at this machine's effective 9.2 tok/s. The ceiling accepted here is 2 hours a day (25% duty on an 8-hour day), because laptops throttle under sustained load.
Effective speed is below comfortable chat streaming (15 tok/s) once prompt-reading is paid for, fine for background work (classification, redaction, summarising), wrong for a chat window someone watches. The routable list is background work on purpose.
What this is not: a serving cluster. Apple runtimes do not batch concurrent users and a fanless laptop sheds ~55% of its package power after 20 minutes of sustained load (Notebookcheck, M4 Air). Each Mac carries its own user, nothing more.
The annual bill, three ways
- All tokens on Claude Haiku 4.5$46,375/yr
The do-nothing baseline.
- Hybrid: 25% on-device, rest on the API$34,781/yr
On-device leg priced at an unmeasured $0 floor, the true number is above this by whatever the electricity turns out to be.
- Per-seat assistant for everyone (M365 Copilot)$180,000/yr
US$30/user/month list, a different product with a different ceiling, shown for scale.
The refresh question
Speccing MacBook Air M5: 16 → 24GB at the next refresh costs $4/device/month over a 4-year cycle, against US$30/user/month for a Copilot seat ($26/month less). They are not substitutes: the RAM buys headroom for on-device routable work, the seat buys a frontier assistant. The point is scale, the entire hardware cost of this story is smaller than one seat licence.
Your workload
Three of these have no published default. Where one is offered, it says where it came from.
170,000 output tokens per person per day. Derived, not observed, nobody publishes tokens per worker, so this is Anthropic's Claude Code docs (US$215 a month, typical) divided by a stated model price.
Two primary measurements disagree by 4×: 15:1 across 100 trillion OpenRouter tokens, 69.7:1 in a Microsoft production trace. This is the dial that decides whether local hardware keeps up, because every generated token waits on the prompt being read first.
Cached input costs a tenth of fresh input. Higher leverage on the bill than model choice.
Not total users, how many are mid-request at the same moment. This binds before volume does on every local option.
Only machines orderable in August 2026. The 512GB and 256GB Mac Studio were withdrawn this year; an M5 Ultra does not exist.
Only classes with a published benchmark on this machine. A pair nobody measured is not modelled.
Where a source published more than one context length, the nearest measurement is used, the same machine and model move from 1,325 tok/s prompt processing and 87.9 generation at 4k to 2,537 and 64.5 at 32k.
The same class of model you would run locally, rented by the token. This is the honest test of "should we self-host", and the one most comparisons skip.
How fast each user's tokens arrive on GPU. Demanding 91 instead of 53 costs 4.8× the throughput on identical silicon, how fast it feels and what it costs are the same dial.
Applied as a veto, not a premium. No defensible dollar value for privacy exists, so none is invented.
No sourced figure exists for either a Mac fleet or a GPU node, so this starts at $0, an explicit zero you are accepting, not a silent one.
How this page works
Four steps, in this order. The answer is usually decided in step two, before cost is even reached.
1. Who is asking, and how much
How many people use AI, what they use it for, and how many are mid-request at the same moment. That last one matters more than the total, and it is the number most comparisons leave out.
2. Can the machine actually do it
Two checks, both before any money is discussed. Does the model physically fit in the machine's memory? And once it is running, is it still fast enough to read comfortably when everyone is using it at once?
3. What each option costs
Five ways to buy the same thing, a cloud API, a per-person licence, hardware on desks, rented GPUs, or a GPU server you own, all reduced to one number: what a million words of output costs.
4. What you are not allowed to do
If your data cannot go to a third party, some options are not expensive, they are unavailable. Those get removed before anything is priced, because a rule is not a line item.
Three words worth knowing
- Tokens: the unit everything is billed in
- Roughly three quarters of a word. A million output tokens is about 750,000 words, a decent novel, or a few weeks of one busy person's AI use.
- Reading vs writing: two different speeds
- A model reads your question before it writes an answer, and the two happen at very different speeds. Most comparisons quote only the writing speed, which is why local hardware looks about twice as capable as it is. (The technical terms are prefill and decode.)
- At the same time: the thing that separates a Mac from a GPU
- A graphics card serves many people at once and gets more efficient as it does. A Mac does not: its total speed stays roughly the same however many people are waiting, so they share it. Four people on one Mac each get a quarter. That single difference decides most of these comparisons.
On these inputs
Cloud API: gpt-oss-120B hosted (median provider)
$2,869 a year · $2.70 per million output tokens · next is Rented H100 capacity at $5,060 (1.8×)
Nobody has measured this machine's draw under inference, so its energy line is missing rather than zero. The figure above is a floor.
Owning the GPU node only beats renting it above 64% sustained utilisation, held for years. This demand uses 12.7% of what the fleet could produce. On these numbers, rent.
Ranked on annual cost among options that survive the memory, concurrency, feasibility and privacy gates.
The two gates
Both are evaluated before cost, and both routinely decide the answer on their own.
1. Does the model fit in memory?
Fits with 30.8GB of headroom.
2. Is it still readable at peak?
Apple aggregate throughput is flat as users are added, measured across MLX, llama.cpp and Ollama, all unchanged from one to four concurrent requests. So users divide one machine's output rather than each getting their own. One machine carries 2 users above the floor, so holding 5 needs 3 of them, that is the quantity priced below, not a disqualification.
Why the second number is lower than the first
Every token generated waits on the prompt being read first, so the rate that matters is 1 / (R / prefill + 1 / decode). At 15:1 on this machine, reading the prompt consumes 50% of its time, so quoting the decode rate alone would overstate capacity by 2.0×. Prompt processing is compute-bound and is where Apple silicon is structurally weakest: on the same 120B MoE, a DGX Spark prefills faster than it decodes relative to Apple hardware, which is why a prompt-heavy workload can favour a machine that looks slower on the headline token rate.
Benchmark: 87.9 tok/s decode, 1,325 tok/s prefill · batch 1 · mlx_lm · 4,096 context · hardware-corner M5 Max local LLM benchmarks
Cost of a million output tokens
1.1bn output + 15.9bn input tokens a yearBars are logarithmic, the spread here is four orders of magnitude and a linear scale would hide it.
- Cloud API: gpt-oss-120B hosted (median provider)$2.70$2,869/yr
US$0.14 in / US$0.6 out per million
- Per-seat: Microsoft 365 Copilot$8.47$9,000/yr
US$30/user/month x 25 people
- MacBook Pro 16" M5 Max (128GB, maxed) x 4$12.74$13,532/yr
4 machines over 3 years at $10,149 each · energy unmeasured on this machine, so this is a floor
- Rented H100 capacity$4.76$5,060/yr
1 GPU held for 2,000 hours a year at US$2.53/hr, you pay to have them available, not only while they generate · your concurrency fills 18% of the batch they could run
- Owned 8x H100 node$107$114,177/yr
Depreciation, power and colocation, this demand uses 12.7% of what the fleet can produce
Break-even against gpt-oss-120B hosted (median provider)
This API already costs less per token than running the model locally. There is nothing to pay back.
Buy or rent the GPUs
64%sustained utilisation
Below this, renting wins. Owning an H100 node also has to clear a second bar it currently fails: rented Blackwell produces tokens for roughly half what an owned H100 does at perfect duty, though Blackwell supply is described as extremely limited, so that rate may not be obtainable at volume.
What this model will not claim
Each of these is a conclusion the evidence did not support, and each one is routinely asserted elsewhere.
Macs replace cloud LLM APIs at scale.
They do not. No published, named-enterprise case study of a Mac fleet replacing cloud LLM APIs survived verification. The one vendor case study found specifies a "Mac Studio M4 Ultra", a chip Apple has never shipped.
AI governance and local inference are the same argument.
Jamf AI Governance controls cloud AI tools running on Macs. It neither requires nor evidences local inference. Conflating them is the fastest way to lose a procurement review.
Regulation mandates on-device inference.
No regulator we found requires it. The closest real forcing function is that HHS OCR treats LLM providers handling PHI as business associates, so a BAA is needed, local inference sidesteps that contracting exercise.
Apple's on-device model can replace an enterprise assistant seat.
The on-device Apple foundation model is roughly 3B parameters with an 8,192-token context window. Private Cloud Compute extends context to 32k. Neither is a Copilot substitute.
Buying an 8x H100 node is a sensible 2026 purchase.
Owned H100 at 100% duty produces tokens at roughly $0.165/M. Rented Blackwell produces them at roughly $0.081/M. H100 is two generations behind the shipping frontier.
Numbers we could not source
Tokens consumed per knowledge worker per day
No credible published data exists. Vendors publish dollars, seats and growth rates. The defaults here are derived from Anthropic's published per-user spend divided by a stated model price, and are labelled as derived wherever they appear.
Sysadmin cost per machine per year
No sourced figure exists for a Mac fleet or a GPU node. The input starts at $0 so the omission is visible rather than silent.
The dollar value of privacy, residency or compliance risk
None is defensible, so privacy removes options rather than pricing them. The single exception is Anthropic's US-residency surcharge of 1.1x on every token category, the only published price for a residency guarantee we found.
Open-weight model API prices (Together, Fireworks, Groq, DeepInfra)
Not researched, and this is the most important missing comparator: hosted open-weight inference is the correct baseline for "should we self-host". Comparing against a frontier API instead favours the local option.
Apple hardware residual value
Not researched. Local hardware is amortised to zero, which understates the Mac case rather than flattering it.
Concurrency benchmarks for a 70B+ model on any Mac
None exists at any batch size above one. Aggregate throughput is assumed flat on the strength of a 9B measurement across four backends, where MLX, llama.cpp and Ollama were all unchanged from one to four concurrent requests. A 0.6B model on an M4 Max did scale 3.7x from 1 to 16, so the flat assumption is drawn from the larger model, and a small enough one may behave differently.
How to read a Mac benchmark
Identical hardware running an identical model spans 3× on token rate depending only on the runtime, MLX 130, llama.cpp 89, Ollama 44 tokens/sec on one measured comparison. A figure without a named tester, runtime version, quantisation, context length and batch size cannot be checked and should be assumed fabricated. A cluster of sites currently publishes precise-looking tables for a “Mac Studio M5 Max” and an “M4 Ultra”, neither of which Apple has ever shipped. Every benchmark on this page names its tester, runtime, quantisation and batch size, and links its source. Where a source did not disclose the context length, the row says so rather than inventing one, and where a source could not be retrieved at all, the benchmark was removed rather than cited.
Compiled 2026-08-06 from vendor pricing pages, published independent benchmarks, EIA electricity statistics, CBRE colocation data and the Uptime Institute data centre survey. Every figure carries a source, a stated derivation, or an admission that it is missing. Two conventions are neither: a 250-day working year and a three-year amortisation for local hardware. Both scale every annual figure on this page and both are choices, not findings.
Where every number comes from
All 53 figures this page can use, each with a link and a label saying what kind of evidence it is. A manufacturer's rating and somebody's actual measurement are not the same thing, so they are never shown as though they were. Last checked 2026-08-06.
- List price
- The published list price. Most buyers negotiate below it.
- Measured
- A third party measured it and published their method.
- Measured once
- One party measured it. Nobody has replicated it yet.
- Reseller price
- A reseller's price, because the vendor does not publish one.
- Our assumption
- A convention we chose, not a finding. Change it if yours differs.
- Official statistic
- A government or official statistical series.
- Rental rate
- A published hourly rental rate.
- Our arithmetic
- Our own arithmetic on sourced inputs. The working is shown.
Validate your numbers with a fleet expert
30 minutes with an Apple fleet consultant: pressure-test the assumptions, map your rollout risks, leave with next steps. Free, no pitch.