How the calculator works
One formula, a table of constants with their sources, and a comparison against every measurement I have. If the model and a meter disagree, the page shows the disagreement.
The formula
Generating one token means reading every active weight and the whole key-value cache from memory once. On the machines compared here, memory bandwidth is the limit, not arithmetic. So the speed ceiling at batch size one is:
tokens per second = efficiency × peak bandwidth ÷ (active weights in GB + KV cache in GB)
- Weights = parameters × bits per weight ÷ 8. A mixture-of-experts model stores all its parameters but reads only the active ones for each token.
- KV cache = 2 × layers × KV heads × head dimension × context length × bytes per value. It uses KV heads, not the hidden size. Using the hidden size overstates the cache two- to eight-fold.
- Energy per token = power at decode ÷ tokens per second. Electricity per million tokens = kilowatts × price per kWh × hours to generate a million tokens.
- Hardware per million tokens = price ÷ (tokens per second × share of time in use × lifetime in seconds) × one million. For a rented GPU it is the hourly rate over tokens per hour.
The page shows a range, not a point. The range is the ceiling times an empirical fraction, taken from the comparison below.
Constants and where they come from
| Machine | Peak GB/s | Efficiency | Power W (dense / mixture) | Basis |
|---|---|---|---|---|
| Unified-memory APU, 256-bit LPDDR5X-8000 | 256.0 | 0.84 | 140 / 171 | Efficiency measured Power measured |
| About 215 GB/s sustained on this class of machine (Framework community forum, 2025-07-22); 83 to 88% in the owner's ledger, 2026-10-10. Smart-plug wall power, one run each, 2026-10-10: 139.6 W (9.7B dense), 170.7 W (31B mixture). Other sizes assumed equal to the nearest class. | ||||
| Desktop, dual-channel DDR5-6000 | 96.0 | 0.75 | 150 / 150 | Efficiency assumed Power assumed |
| No measurement found. A plausibility range.. Assumed whole-system draw with a CPU under inference load. | ||||
| Server, 8-channel DDR5-4800 | 307.2 | 0.75 | 380 / 380 | Efficiency assumed Power assumed |
| STREAM runs on larger server parts reach about 82% of peak; this part was not measured.. Assumed whole-system draw. | ||||
| Data-centre GPU, HBM3 (H100 SXM class) | 3350.0 | 0.75 | 315 / 315 | Efficiency assumed Power assumed |
| No measurement found.. 45% of the 700 W board rating. Published decode measurements report 34 to 60% of rated power (see sources). GPU board only: excludes host, networking and cooling. | ||||
Does the model match what was measured?
Each row compares the speed ceiling with a real run on the unified-memory machine. A ratio near 1 means the ceiling is tight. Below 1 means the machine runs slower than the bandwidth limit.
| Run | GB read per token | Ceiling tok/s | Observed tok/s | Ratio | Evidence |
|---|---|---|---|---|---|
| 9.7B dense, Q4_K_M | 6.14 | 35.0 | 36.68 | 1.05 | Measured Smart-plug run, 2026-10-10, one run |
| 31.1B mixture, 3B active, Q4_K_M | 1.76 | 122.2 | 67.32 | 0.55 | Measured Smart-plug run, 2026-10-10, one run |
| 70B dense, Q4_K_M | 43.32 | 5.0 | 4.5 to 5.1 | 0.91 to 1.03 | Measured Published llama-bench, 2025-07-22 and repo to 2026-03 |
| 32B dense, Q4_K_M | 20.18 | 10.7 | 11.3 | 1.06 | Measured Published llama-bench, repo to 2026-03 |
| 14B dense, Q4_K_M | 9.15 | 23.5 | 24.5 | 1.04 | Measured Published llama-bench, repo to 2026-03 |
| 30B mixture of experts, 3.3B active, Q4_K_M | 2.07 | 104.0 | 66.3 to 86.1 | 0.64 to 0.83 | Measured Published llama-bench, 2025-07-22 and repo to 2026-03 |
Dense models land at the ceiling. Mixture models run well below it, because routing between experts and small reads cost more than the bandwidth sum suggests. The calculator turns this into an expected range of 0.9 to 1.07 of the ceiling for dense models and 0.55 to 0.85 for mixtures. The build fails if a row above falls outside its range.
A third metered run, on a 125B mixture model, is excluded. Its recorded speed implies 356 GB/s of memory traffic on a 256 GB/s bus. Either its speed or its active-parameter count is wrong in the ledger.
What it does not model
- Batching. Serving many requests at once reuses each weight read many times. It changes energy and cost per token several-fold and is the main reason data-centre figures differ from this page.
- Prompt processing. The first tokens are limited by arithmetic, not bandwidth, and are not estimated here.
- Other machines. Only the unified-memory machine is validated. Efficiency and power for the other three are assumptions, marked as such.
- Long contexts. Attention compute grows with context and is ignored. The KV read is included.
- Prices. Electricity, hardware and API prices move. The defaults are dated and editable.
Sources
- Peak bandwidth of a 256-bit LPDDR5X-8000 bus is 256 GB/s. AMD Ryzen AI Max+ 395 product page, fetched 2026-10-11, and the arithmetic 8000 MT/s x 256 bit / 8.
- About 215 GB/s sustained on this class of APU. Framework community forum, Strix Halo llama-bench tests, 2025-07-22.
- H100 SXM: 3.35 TB/s HBM3, 80 GB, up to 700 W. NVIDIA H100 page, fetched 2026-10-11.
- An H100 uses about 34% of its rated power during decode. arXiv:2602.18568, "RPU: A Reasoning Processing Unit", 2026-02-20, verified in the full text 2026-10-11.
- NVIDIA GPUs use 45 to 60% of rated power during decode. arXiv:2604.10852, "The xPU-athalon", 2026-04, verified in the full text 2026-10-11.
- Bits per weight of llama.cpp quantisations. llama.cpp quantize README.
- US electricity prices, July 2026. EIA Electric Power Monthly, release 2026-09-24.
- H100 rental about $3 per GPU-hour. Akash Network H100 rental price overview, August 2026.
- A 30B-A3B-class open-weight model near $0.50 per million output tokens. DeepInfra model page, fetched 2026-10-11, undated.
- EIA Electric Power Monthly, release 2026-09-24: US residential 18.31 cents per kWh, commercial 14.53 cents per kWh, July 2026.
- A 30B-A3B-class open-weight model was listed at $0.50 per million output tokens by one provider (DeepInfra, undated page, fetched 2026-10-11). Prices vary widely. Set your own.
Model sizes and attention layouts come from public configuration files of open models, described here by size and kind.
The formula is implemented once in model.js and checked against a separate Python implementation on 400 random inputs, with zero disagreements at build time.