Qvunex
What your GPUs were billed, against what they actually served.
A GPU hour has one price. What it delivers does not. Qvunex measures self-hosted inference cost from the server's own counters: cold starts, idle time, cost per request and per million tokens.
97%of one H100 cold wake was boot, before the server could answer
$1.26 vs $0.22per 1M tokens, same H100 hour, 1 request vs 8 at once
5.7xoverstatement when cost is priced from one request's speed
Measured on vLLM 0.27.1, rented H100 SXM and T4, September 2026. Conditions and raw data are public.
What we measure
- Billed against served: every paid window marked served, paid with nothing served, or unknown. Never a guess.
- Cold starts: boot, serving and idle tail of each wake, in seconds and dollars. On SGLang the health check said ready while the first request still took 176 seconds more.
- Cost by load: whole-server output at the load you actually run, on vLLM and SGLang.
- Conditions: GPU, host CPU, versions and settings written down, so anyone can repeat it.
Work with us
- Measured report for GPU clouds and hardware teams: your GPUs, your prices, published as your own data. From $1,500.
- Cost audit for teams paying the bill: where the money on your GPUs goes. $1,200, paid only after you read the findings.
Open method
The method is public: How to price one GPU wake. The tool is one open-source file with its own self-test: github.com/qaisermehdi3-coder/qvunex.