← r/LocalLLaMA
▲
0
 
3👁
r/LocalLLaMA · u/Distinct-Pie2389 · 5h ago

llm-tune: getting local models that actually perform

post image

Made a little agent skill called llm-tune to help find the best settings for running local LLMs on your hardware.

Still working on it, but I'm looking to test it across more GPU setups. If you try it out, feedback and benchmark results are welcome.

Supported:

  • Architectures: Dense, MoE, hybrid MoE/Mamba
  • GPUs: NVIDIA, AMD, Intel Arc
  • Apple Silicon: M-series Macs via MLX
  • Engines: llama.cpp, Ollama, vLLM
  • Tuning: Quantization, context/KV cache, GPU offloading, MTP, sampling, reasoning, and agent harness settings
  • Benchmarks: VRAM/RAM usage, tokens/sec, context recall, and output quality

Currently measured on an RTX 4090; other hardware and backends are documented but need more real-world testing.

1 0 0 10/9 03:33 10/9 07:36 UTC
scorecomments3 sightings
first seen 2026-10-09 03:33 UTClast seen 2026-10-09 07:36 UTCscore then 0score now 0gained 0sightings 3
open on reddit ↗ 💬 7 (+3)