← r/LocalLLaMA
▲
0
 
1👁
r/LocalLLaMA · u/paq85 · 4h ago

How much does your local LLM server really cost? (power draw + tok/s, self-hosted vs cloud API, free in-browser) :)

I've been tuning local LLM setup for many months, and the number people actually want is the electricity bill. This calculator takes your GPU's power draw (W) and token speed (tok/s for prompt and decode), plugs in your energy price per kWh, and shows cost per hour, per request, monthly, and cost per 1M input tokens — with and without KV cache.

There are 3 GPU presets (RTX 4070 Ti Super, RTX 5090, RTX 5090 eco) if you want to start from realistic numbers and adjust from there. The math treats cached tokens at 1/100 the time of uncached, so the realistic scenario is fixed at 60k input / 50k cached / 2k output / 90% uptime.

The cloud API comparison compares each self-hosted profile against a reference API ($0.25/1M input, $1.20/1M output) at the same scenario, so you can see whether self-hosting actually saves money or costs more.

Everything runs in your browser, nothing is uploaded and there is no sign-up. From my experience, the per-request number is the one worth comparing with the cloud — the monthly bill is just the per-request cost times your actual request rate.

https://appdoesit.com/apps/llm-cost-calculator — it's one of the 139 free tools in the catalog.

Try it by yourself :) — how does your server compare?

posted Fri, 09 Oct 2026 08:39:14 GMTseen 1 time
open on reddit ↗ 💬 20