Release: Qwen-2B-RCOL Dynamic Low-Bit Quantization (IQ1_M, IQ2_M, IQ3_M)
Hey guys - primarily a research release with working models,
Not so much a model as a psuedo-new quantization technique. I've been experimenting with a modification of ISTALab's RCO algorithm that can quantize models on a strict VRAM budget.
It's at its core an approximation algorithm that attempts to make up the difference with a few different strategies. I'm quite happy with how well it's working at low bits. I've seen significant KLD improvements, particularly in the 1 bit and 2 bit range. This should be model agnostic and it doesn't require loading the FP16 teacher into VRAM. I've included a little more detail on the model card and will probably publish the full recipes. I plan to apply this method to the larger models in the family next.
Unfortunately Unsloth does not have their KL divergence table available, but I've created a table for you to see the difference between these quants and standard llama.cpp imatrix quants. For accuracies sake, both my quants and the llama.cpp quants are calibrated on the sane wikitext imatrix dataset, and tested on two held out sets.
As you can see, particularly at extremely low BPW, these quants vastly outperform their standard imatrix KLD - particularly drastically cutting IQ1\_M divergence at nearly the same size budget.
These are more proof of concept quants, as 2B is a fairly small model and suffers from such low BPW, but here's a fun example of how these are still relatively functional at extreme compression
and here's the standardized imatrix IQ1\_M quant (not using dynamic RCOL)
and the file sizes
GGUFs: