How to know about optimized engines
Optimizing an engine for a family of models and hardware combination seems very appealing. As someone who uses qwen3.6 and 3.8 a lot, and is evaluating hardware options before going fully local, it's hard to keep up with the state of things.
Huggingface made it possible to see the development branches and derivative modifications to models. Is there something similar for inference engines yet? I've seen some where the it's optimized for a shell game of moving layers between vram and ram while using ngrams (amazing), others are all about quants (less amazing), but it's incredibly difficult to compare apples to apples where there's variability on card architecture, vram size, quant approach, memory management optimization approach... I was already busy over thinking my vram selection, now it's an even bigger decision matrix without any filters!