Qvunex

What your GPUs were billed, against what they actually served.

A GPU hour has one price. What it delivers does not. Qvunex measures self-hosted inference cost from the server's own counters: cold starts, idle time, cost per request and per million tokens.

97%of one H100 cold wake was boot, before the server could answer
$1.26 vs $0.22per 1M tokens, same H100 hour, 1 request vs 8 at once
5.7xoverstatement when cost is priced from one request's speed

Measured on vLLM 0.27.1, rented H100 SXM and T4, September 2026. Conditions and raw data are public.

What we measure

Work with us

See a sample report

Open method

The method is public: How to price one GPU wake. The tool is one open-source file with its own self-test: github.com/qaisermehdi3-coder/qvunex.