QuantBench

Independent 4-bit quantization benchmarks for small instruct models

Independence

No vendor, lab, or model author paid for placement on this site, reviewed the results before publication, or influenced how the rows are ordered or presented. There is no sponsored row, no paid tier, and no advertising.

This benchmark was run by one independent researcher (Faisal), working solo, on a personal Modal account and on personal time. No employer resources, data, or hardware were used. The compute was Modal cloud GPUs — A10 and T4, preemptible instances — paid for out of expiring Modal credit on that personal account; the driver's own estimate for the surviving run is about $27. Only permissively licensed public models were tested.

Results are reproducible from the pinned configuration published in the methodology: quantization parameters, calibration sets and sizes, seeds, perplexity and throughput protocols, and the torch / transformers / quantization-library versions recorded per row. Failed runs are published with their error strings rather than dropped.

Corrections are welcome. If a row does not reproduce, report it and it will be marked or withdrawn.