How much does your local LLM server really cost? (power draw + tok/s, self-hosted vs cloud API, free in-browser) :)
I've been tuning local LLM setup for many months, and the number people actually want is the electricity bill. This calculator takes your GPU's power draw (W) and token speed (tok/s for prompt and decode), plugs in your energy price per kWh, and shows cost per hour, per request, monthly, and cost per 1M input tokens — with and without KV cache.
There are 3 GPU presets (RTX 4070 Ti Super, RTX 5090, RTX 5090 eco) if you want to start from realistic numbers and adjust from there. The math treats cached tokens at 1/100 the time of uncached, so the realistic scenario is fixed at 60k input / 50k cached / 2k output / 90% uptime.
The cloud API comparison compares each self-hosted profile against a reference API ($0.25/1M input, $1.20/1M output) at the same scenario, so you can see whether self-hosting actually saves money or costs more.
Everything runs in your browser, nothing is uploaded and there is no sign-up. From my experience, the per-request number is the one worth comparing with the cloud — the monthly bill is just the per-request cost times your actual request rate.
https://appdoesit.com/apps/llm-cost-calculator — it's one of the 139 free tools in the catalog.
Try it by yourself :) — how does your server compare?