What does a model cost to run?
Pick a model and a machine. Memory, decode speed, energy per token and cost per million tokens appear with their assumptions beside them.
Decode speedModelled
–
Energy per tokenModelled
–
Cost per million tokensModelled
–
Memory
Decode speed, tokens per second
Cost per million tokens
Metered on this machine Measured
| Model, one run | tok/s | Wall W | J/token |
|---|---|---|---|
| 9.7B dense, Q4_K_M | 36.7 | 140 | 3.81 |
| 31.1B mixture, 3B active, Q4_K_M | 67.3 | 171 | 2.54 |
Smart-plug wall power averaged over a single run of 6 to 10 seconds, 2026-10-10 23:26 UTC. Includes the whole machine. The record and its limits, and the results file.