QuantBench

Independent 4-bit quantization benchmarks for small instruct models

How to read this. Lower Δ PPL vs fp16 is better (quality lost to quantization, measured on wikitext2-test). tok/s is greedy, batch size 1, 64→256 tokens, median of 3, prefill included — it is a comparable number across rows here, not a claim about optimized serving throughput. Each fp16 baseline is stack-specific, so deltas are always computed against the baseline from the same quantization stack. Full protocol on the methodology page.

ModelMethodGPUCalibrationSeedΔ PPL vs fp16tok/sVRAM (GB)Artifact (GB)BackendQuant lib
16 configurations that failed (kept on purpose)
Failures are results too: they say which stack/model combinations do not work today, and why. They are excluded from the leaderboard above.
ModelMethodGPUCalibrationSeedError
deepgrove/BonsaiawqA10wikitext2-train n=1280OSError("/cache/quants/bonsai-0p5b__awq__wikitext2__n128__s0 does not appear to have a file named modeling_qllama.py. Checkout 'https://huggingface.co//cache/quants/bonsai-0p5b__awq__wikitext2__n128__s0/tree/main' for available files.")
deepgrove/BonsaiawqT4wikitext2-train n=1280OSError("/cache/quants/bonsai-0p5b__awq__wikitext2__n128__s0 does not appear to have a file named modeling_qllama.py. Checkout 'https://huggingface.co//cache/quants/bonsai-0p5b__awq__wikitext2__n128__s0/tree/main' for available files.")
deepgrove/BonsaiawqA10wikitext2-train n=1281OSError("/cache/quants/bonsai-0p5b__awq__wikitext2__n128__s1 does not appear to have a file named modeling_qllama.py. Checkout 'https://huggingface.co//cache/quants/bonsai-0p5b__awq__wikitext2__n128__s1/tree/main' for available files.")
deepgrove/BonsaiawqT4wikitext2-train n=1281OSError("/cache/quants/bonsai-0p5b__awq__wikitext2__n128__s1 does not appear to have a file named modeling_qllama.py. Checkout 'https://huggingface.co//cache/quants/bonsai-0p5b__awq__wikitext2__n128__s1/tree/main' for available files.")
deepgrove/BonsaiawqA10wikitext2-train n=1282OSError("/cache/quants/bonsai-0p5b__awq__wikitext2__n128__s2 does not appear to have a file named modeling_qllama.py. Checkout 'https://huggingface.co//cache/quants/bonsai-0p5b__awq__wikitext2__n128__s2/tree/main' for available files.")
deepgrove/BonsaiawqT4wikitext2-train n=1282OSError("/cache/quants/bonsai-0p5b__awq__wikitext2__n128__s2 does not appear to have a file named modeling_qllama.py. Checkout 'https://huggingface.co//cache/quants/bonsai-0p5b__awq__wikitext2__n128__s2/tree/main' for available files.")
deepgrove/Bonsainone-fp16A10none n=00a10 batch call failed: RuntimeError("ImportError: cannot import name 'LossKwargs' from 'transformers.utils' (/usr/local/lib/python3.11/site-packages/transformers/utils/__init__.py)")
deepgrove/BonsaigptqA10wikitext2-train n=1280a10 batch call failed: RuntimeError("ImportError: cannot import name 'LossKwargs' from 'transformers.utils' (/usr/local/lib/python3.11/site-packages/transformers/utils/__init__.py)")
deepgrove/BonsaigptqA10wikitext2-train n=1281a10 batch call failed: RuntimeError("ImportError: cannot import name 'LossKwargs' from 'transformers.utils' (/usr/local/lib/python3.11/site-packages/transformers/utils/__init__.py)")
deepgrove/BonsaigptqA10wikitext2-train n=1282a10 batch call failed: RuntimeError("ImportError: cannot import name 'LossKwargs' from 'transformers.utils' (/usr/local/lib/python3.11/site-packages/transformers/utils/__init__.py)")
Qwen/Qwen2.5-3B-InstructawqA10openhermes-2.5 n=5120OutOfMemoryError('CUDA out of memory. Tried to allocate 4.96 GiB. GPU 0 has a total capacity of 22.06 GiB of which 419.44 MiB is free. Process 1 has 21.64 GiB memory in use. Of the allocated memory 16.42 GiB is allocated by PyTorch, and 4.92 GiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid
Qwen/Qwen2.5-3B-InstructawqA10openhermes-2.5 n=5121OutOfMemoryError('CUDA out of memory. Tried to allocate 4.87 GiB. GPU 0 has a total capacity of 22.06 GiB of which 4.56 GiB is free. Process 1 has 17.49 GiB memory in use. Of the allocated memory 16.76 GiB is allocated by PyTorch, and 436.95 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid
Qwen/Qwen2.5-3B-InstructawqA10openhermes-2.5 n=5122OutOfMemoryError('CUDA out of memory. Tried to allocate 4.91 GiB. GPU 0 has a total capacity of 22.06 GiB of which 3.35 GiB is free. Process 1 has 18.70 GiB memory in use. Of the allocated memory 16.86 GiB is allocated by PyTorch, and 1.54 GiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fr
Qwen/Qwen2.5-3B-InstructawqA10wikitext2-train n=5120OutOfMemoryError('CUDA out of memory. Tried to allocate 4.72 GiB. GPU 0 has a total capacity of 22.06 GiB of which 1.27 GiB is free. Process 1 has 20.78 GiB memory in use. Of the allocated memory 15.79 GiB is allocated by PyTorch, and 4.69 GiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fr
Qwen/Qwen2.5-3B-InstructawqA10wikitext2-train n=5121OutOfMemoryError('CUDA out of memory. Tried to allocate 4.71 GiB. GPU 0 has a total capacity of 22.06 GiB of which 427.44 MiB is free. Process 1 has 21.63 GiB memory in use. Of the allocated memory 16.48 GiB is allocated by PyTorch, and 4.85 GiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid
Qwen/Qwen2.5-3B-InstructawqA10wikitext2-train n=5122OutOfMemoryError('CUDA out of memory. Tried to allocate 4.62 GiB. GPU 0 has a total capacity of 22.06 GiB of which 411.44 MiB is free. Process 1 has 21.65 GiB memory in use. Of the allocated memory 17.10 GiB is allocated by PyTorch, and 4.25 GiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid