← r/LocalLLaMA
▲
25
+16
19👁
r/LocalLLaMA · u/tom_tsai28 · 5d ago

[Discussion] A 5KB pure x86-64 assembly engine for Gemma-2B (FP16, 4.6 tok/s on CPU)

Hi everyone,

Sharing a personal project exploring the minimal bare-metal footprint required to run an autoregressive LLM.

Instead of relying on large runtimes or compiler abstractions, I wrote an inference engine for Gemma-2B entirely in flat x86-64 assembly (FASM):

\- \*\*Binary footprint\*\*: Total 5.2 KB flat machine code (\gemma\_engine.bin\ 3.7 KB + \mat\_smp\_f16c\_gemm\_avx2.bin\ 1.5 KB).

\- \*\*Execution\*\*: Pure AVX2 + F16C with custom 4-thread SMP GEMM for prefill. Sustains \~18.5 GB/s memory bandwidth on commodity DDR4-2400.

\- \*\*Decoding\*\*: 4.5 \~ 4.7 tokens/s in FP16 on an older quad-core i5 desktop.

\- \*\*Dependencies\*\*: Zero C/C++ runtime, zero PyTorch. The Python harness only uses \ctypes\ for \VirtualAlloc\ and OS threads.

This isn't meant to compete with feature-complete tools like llama.cpp. Rather, it's a first-principles exploration to see how cleanly a modern Transformer can be mapped to raw silicon, and to serve as a reference point for future micro-LLMs on resource-constrained microcontrollers (MCU/DSP).

The repository is open source:

\- GitHub: https://github.com/tomtsai28/PULSAR-ASM

\- Architecture notes: https://github.com/tomtsai28/PULSAR-ASM/blob/main/doc/pulsar\_asm\_cpu\_limit\_retrospective.md

Any code audits, observations, or thoughts on bare-metal inference are welcome.

26 0 25 10/4 05:28 10/7 17:48 UTC
scorecomments19 sightings
first seen 2026-10-04 05:28 UTClast seen 2026-10-07 17:48 UTCscore then 9score now 25gained +16sightings 19
open on reddit ↗ 💬 19 (+15)