← r/LocalLLaMA
▲
11
+8
33👁
r/LocalLLaMA · u/hiImMate · 6d ago

gufo_windows pre-package for Strix Halo users

In the latest release of the completely unofficial gufo port to windows I've added a pre-packaged library that you can use to try out gufo for yourself. No need to build anything just grab the .zip and unpack it.

I've also added start.cmd for easily starting the server, it will try to autodiscover supported quants for easy startup.

current support on windows:
3.8 Flash Next: UD\_Q4\_XL

27b: UD\_Q4\_XL

35BA3B: UD\_Q8\_XL + TeilCoder (I assume ornith as well since its the same but untested).

Any issues you run into please submit an issue to github or here.

I am mainly making this for myself but happy to share as I only run gufo with Flash Next now. It is solid 40tps avg on agentic even at higher ctx.

Important to set your VRAM to 96gb! Although its unified, windows adds overhead for reading 'shared' ram vs 'dedicated' vram.

other AMD users: I'm sorry but the library is specifically for gfx1151, I don't have any other card, therefore I can't check or add support to anything else.

psa: yes this is vibecoded, I run a logit check and the model's output must stay bit-identical after changes.

11 0 11 10/3 08:30 10/9 04:18 UTC
scorecomments33 sightings
first seen 2026-10-03 08:30 UTClast seen 2026-10-09 04:18 UTCscore then 3score now 11gained +8sightings 33
open on reddit ↗ 💬 14 (+13)