QuantBench

Independent 4-bit quantization benchmarks for small instruct models

Inference Cost & Quantization Audit

$3,000 · 5 business days · fully async · no meetings required

You are paying a GPU bill you suspect is 2–3× larger than it needs to be, and the advice you can find is either a vendor pitch or a benchmark run on hardware you do not rent. This audit answers one question with measurements on your model and your GPU tier: what do you actually give up, and what do you actually save, at each quantization level?

What you get

1. A quantization trade-off table for your model(s). Quality delta versus the FP16 baseline, tokens/sec, peak VRAM, and $/1k queries at each level (GPTQ / AWQ / INT8, and FP8 where your hardware supports it) — the same format and the same honesty rules as the public leaderboard. Every number is measured, never extrapolated from a model of a different size.

2. A prioritized fix list, matched to your GPU tier, working down the stack in the order that actually pays: serving stack → weight quantization → prefix caching → KV-cache configuration → speculative decoding. Each item carries an expected gain and an honest effort estimate, and items that will not help you are named as such rather than padded in.

3. One reproduction script per recommendation. You re-run it yourself and get the same number, or the recommendation does not ship.

How it runs

Send your model ID (or architecture and size, if it is private), your GPU tier, and a rough traffic profile. Everything after that is asynchronous — you get the deliverable in five business days. No discovery calls, no workshops, no slide deck.

What this is not

deliverable never exceed the measured table.

well-served — it happens, particularly with small batch sizes on modern serving stacks — the deliverable says so, and says it in the first paragraph rather than burying it.

($2,000–$3,000/mo) but are a separate conversation, and capacity is two.

Why trust the numbers

The public QuantBench leaderboard is the same methodology applied to public models: 161 measured rows, the 7 failures published with their error strings, a methodology page carrying 13 numbered limitations, and the exact quantized weights downloadable so anyone can check a row against the artifact that produced it. That benchmark also caught a measurement trap most comparisons miss: the quantization toolchain itself moved measured FP16 throughput by up to 27.8% on identical weights and identical hardware — which means a speedup quoted against someone else's baseline can be mostly environment. Your audit controls for that.

See a worked example built from the public benchmark data — the same structure your deliverable arrives in.

Getting started

Email the address on the independence page with subject AUDIT and include: model ID or architecture + size, GPU tier(s) you run or are considering, rough queries/day and typical prompt/response length, and your deadline. You will get a yes/no on fit within one business day — including a "no, this will not help you" when that is the honest answer.

Get the inference cost playbook

The method behind the audit: the five levers in the order that actually pays, with what our own measurements say about each. Leave an email and the link appears below — no drip sequence, no newsletter, and the address is used only to tell you when the benchmark adds models.

Prefer not to share an email? The playbook is also linked from the navigation — it is a courtesy gate, not a paywall.