4-bit Qwen2.5 that stays closer to fp16 than the official AWQ, on the same vLLM kernel (1.5B and 7B, code + models)
I'm an undergrad. Over the last two weeks I built a quantizer on my MacBook, using Claude as a coding assistant. The results were then reproduced on an NVIDIA A10G by M. Federico (a family member who works in ML), using separate evaluation scripts.
It is GPTQ with three additions: each group's grid is fitted to its weights instead of using min-max, a second pass re-checks every rounded weight, and the grid is refitted against the layer's input statistics. Offsets are integer zero points, so the model packs into the normal AWQ format and runs on vLLM's awq_marlin kernel.
A10G, vLLM 0.29, everything served through the int4 kernel. WikiText-2 perplexity / HumanEval pass@1:
| model | fp16 | mine, 4-bit | official Qwen AWQ 4-bit |
|:--|:--|:--|:--|
| Qwen2.5-1.5B-Instruct | 9.37 / 37.2% | 9.66 / 33.5% | 10.16 / 34.1% |
| Qwen2.5-7B-Instruct | 7.15 / 70.1% | 7.29 / 67.1% | 7.58 / 64.6% |
What this does not show:
- One run per row. The HumanEval differences between the 4-bit models are within noise (about 3.6 points).
- I calibrate on WikiText-2 train, which helps on the perplexity test. On the 1.5B model that was worth about 0.3.
- The lead shrinks as the model gets bigger.
- At 3 bits the method keeps perplexity close but loses more than half of code and math ability. I would not use those for code.
Code and all results, including what did not work: https://github.com/dfed25/mlx-gptq
7B: https://huggingface.co/dfed24/Qwen2.5-7B-Instruct-gptq-4bit-int4-awq
1.5B: https://huggingface.co/dfed24/Qwen2.5-1.5B-Instruct-gptq-4bit-int4-awq
MLX versions: https://huggingface.co/dfed24
vllm serve dfed24/Qwen2.5-7B-Instruct-gptq-4bit-int4-awq --quantization awq_marlin --dtype float16
If anyone tests it on a benchmark I haven't run, I'd like to see the numbers either way.