If you have a 3090, or other 30xx for local LLMs, I have something for you
I have a custom fork of llama.cpp designed around the ampere architecture specifically (though many of the upgrades also translate to faster performance of blackwell + lovelace). The recommended config supports 90+ TPS (for agentic/coding, at temp 1; greedy will of course be faster) through 100K tokens, with context of up to 240K.
If you want the repo, it is here:
https://github.com/JakeATX/llamAmpere
I recommend running with this quant, which is \~ 4 K M quality but considerably faster (technically, a 3 K XL upgrade)
https://huggingface.co/jakeatx/Qwen3.8-27B-ATX-IQ4\_XS-M-GGUF
If you want the deep dive on how it is so much faster (80% vs the near comp at 200K!), at more context, there is a long form article here.
https://x.com/JakeKAllDay/status/2095646450138874095?s=20
Running faster than API speeds on my 3090 (for 27b at least) has genuinely been a step change in the utility of the model + card for me. I hope you enjoy it!