How the calculator works

One formula, a table of constants with their sources, and a comparison against every measurement I have. If the model and a meter disagree, the page shows the disagreement.

The formula

Generating one token means reading every active weight and the whole key-value cache from memory once. On the machines compared here, memory bandwidth is the limit, not arithmetic. So the speed ceiling at batch size one is:

tokens per second = efficiency × peak bandwidth ÷ (active weights in GB + KV cache in GB)

  • Weights = parameters × bits per weight ÷ 8. A mixture-of-experts model stores all its parameters but reads only the active ones for each token.
  • KV cache = 2 × layers × KV heads × head dimension × context length × bytes per value. It uses KV heads, not the hidden size. Using the hidden size overstates the cache two- to eight-fold.
  • Energy per token = power at decode ÷ tokens per second. Electricity per million tokens = kilowatts × price per kWh × hours to generate a million tokens.
  • Hardware per million tokens = price ÷ (tokens per second × share of time in use × lifetime in seconds) × one million. For a rented GPU it is the hourly rate over tokens per hour.

The page shows a range, not a point. The range is the ceiling times an empirical fraction, taken from the comparison below.

Constants and where they come from

Defaults. Every one can be changed in the calculator.
MachinePeak GB/sEfficiencyPower W (dense / mixture)Basis
Unified-memory APU, 256-bit LPDDR5X-8000256.00.84140 / 171Efficiency measured Power measured
About 215 GB/s sustained on this class of machine (Framework community forum, 2025-07-22); 83 to 88% in the owner's ledger, 2026-10-10. Smart-plug wall power, one run each, 2026-10-10: 139.6 W (9.7B dense), 170.7 W (31B mixture). Other sizes assumed equal to the nearest class.
Desktop, dual-channel DDR5-600096.00.75150 / 150Efficiency assumed Power assumed
No measurement found. A plausibility range.. Assumed whole-system draw with a CPU under inference load.
Server, 8-channel DDR5-4800307.20.75380 / 380Efficiency assumed Power assumed
STREAM runs on larger server parts reach about 82% of peak; this part was not measured.. Assumed whole-system draw.
Data-centre GPU, HBM3 (H100 SXM class)3350.00.75315 / 315Efficiency assumed Power assumed
No measurement found.. 45% of the 700 W board rating. Published decode measurements report 34 to 60% of rated power (see sources). GPU board only: excludes host, networking and cooling.

Does the model match what was measured?

Each row compares the speed ceiling with a real run on the unified-memory machine. A ratio near 1 means the ceiling is tight. Below 1 means the machine runs slower than the bandwidth limit.

Efficiency 0.84 × 256 GB/s = 215 GB/s. Context for published rows is not stated by their authors, so I assumed 512 tokens.
RunGB read per tokenCeiling tok/sObserved tok/sRatioEvidence
9.7B dense, Q4_K_M6.1435.036.681.05Measured Smart-plug run, 2026-10-10, one run
31.1B mixture, 3B active, Q4_K_M1.76122.267.320.55Measured Smart-plug run, 2026-10-10, one run
70B dense, Q4_K_M43.325.04.5 to 5.10.91 to 1.03Measured Published llama-bench, 2025-07-22 and repo to 2026-03
32B dense, Q4_K_M20.1810.711.31.06Measured Published llama-bench, repo to 2026-03
14B dense, Q4_K_M9.1523.524.51.04Measured Published llama-bench, repo to 2026-03
30B mixture of experts, 3.3B active, Q4_K_M2.07104.066.3 to 86.10.64 to 0.83Measured Published llama-bench, 2025-07-22 and repo to 2026-03

Dense models land at the ceiling. Mixture models run well below it, because routing between experts and small reads cost more than the bandwidth sum suggests. The calculator turns this into an expected range of 0.9 to 1.07 of the ceiling for dense models and 0.55 to 0.85 for mixtures. The build fails if a row above falls outside its range.

A third metered run, on a 125B mixture model, is excluded. Its recorded speed implies 356 GB/s of memory traffic on a 256 GB/s bus. Either its speed or its active-parameter count is wrong in the ledger.

What it does not model

  • Batching. Serving many requests at once reuses each weight read many times. It changes energy and cost per token several-fold and is the main reason data-centre figures differ from this page.
  • Prompt processing. The first tokens are limited by arithmetic, not bandwidth, and are not estimated here.
  • Other machines. Only the unified-memory machine is validated. Efficiency and power for the other three are assumptions, marked as such.
  • Long contexts. Attention compute grows with context and is ignored. The KV read is included.
  • Prices. Electricity, hardware and API prices move. The defaults are dated and editable.

Sources

Model sizes and attention layouts come from public configuration files of open models, described here by size and kind.

The formula is implemented once in model.js and checked against a separate Python implementation on 400 random inputs, with zero disagreements at build time.