← r/LocalLLaMA
▲
0
 
15👁
r/LocalLLaMA · u/giveen · 5d ago

Welcome to Spite

Spite is a vision I had. What if you could take all those custom inference engines out there, designed for specific cards or setups, and compact them into one system? You get to design the kernels and optimizations for your setup. You only compile for your cards and the models you like to run.

Spite is built on a single rule: every layer is replaceable without touching any other layer.

That sounds abstract, so here's what it means in practice:

\### Every model is its own module

Kernels are grouped by family and variant: \kernels/llama/llama4/\, \kernels/deepseek/v4/\, \kernels/qwen/qwen3\_5/\, \kernels/mistral/mistral4/\, \kernels/gemma/gemma4/\. Adding a new model variant means adding a new \<family>/<model>/\ folder. Nothing about the existing models changes. The dispatcher finds it automatically.

\### Every GPU is its own module

\kernels/llama/llama4/sm\_89/\ is completely separate from \kernels/llama/llama4/rdna3/\. An RTX 4090 kernel can use FP8 tensor cores. An RX 7900 XTX kernel can exploit 96 MB of Infinity Cache. An Apple M4 kernel can use the Neural Engine. Each gets what makes it fast, not a watered-down kernel that has to work on everything.

\### Every operation is independently tunable

Kernels don't have to implement everything. A kernel that only optimizes attention leaves FFN and \rms\_norm\ to the fallback. You tune the one op that's your bottleneck. Later, someone else improves FFN. Both improvements stack automatically—the dispatcher picks the best available kernel for each op on each GPU.

\### Every subsystem is swappable

The sampler, tokenizer, KV cache backend, and offload policy are all plugin registries. Register a custom sampler for a specific model or task, and the engine uses it. Register a custom KV cache for a memory-constrained deployment, and the scheduler uses it. Nothing needs to be forked.

\\\`rust

let engine = EngineBuilder::new()

.with\_sampler(PluginKey::for\_model("llama4"), Box::new(MyGreedySampler))

.with\_cache(PluginKey::default(), Box::new(PagedKvCache::new(vram)))

.build(ExecutorConfig::default());

\\\`

\### Every component is usable standalone

Spite is a Rust workspace. You can use just the loader, just the scheduler, or just the ABI types for kernel development—without pulling in the full server stack. Build what you need from the pieces that fit.

I'm still in very early stages, but I would love people to contribute.

https://github.com/giveen/spite

1 0 0 10/4 23:34 10/8 01:54 UTC
scorecomments15 sightings
first seen 2026-10-04 23:34 UTClast seen 2026-10-08 01:54 UTCscore then 0score now 0gained 0sightings 15
open on reddit ↗ 💬 24 (+8)