← r/LocalLLaMA
▲
16
 
1👁
r/LocalLLaMA · u/Chromix_ · 6h ago

A 24KB stand-alone HTML-LLM that can generate consistent stories

https://preview.redd.it/7sknv6omnguh1.png?width=1023&format=png&auto=…

Want to try? Just go here and click "New seed" a lot. You will get different stories, it just requires some clicking - lucky RNG draws: https://output.jsbin.com/nikupemuta/1
You should even get above 60 tokens per second with a smartphone.

FAQ:

  • Why?! Because it's possible, and fun. My original idea was to simply mash MacroStory and LittleBit together, but that did not work at all.
  • Is it relevant? Not a tiny bit! Even a 3 MB HTML file with external dependencies would be loaded in a second and could serve a way more capable model with a lot less work. Squeezing the model itself to 20 KB at 0.01 KLD was trivial. Saving another 5 KB while maintaining output quality required a lot of time.
  • How? Local Qwen3.8, some ideas, lots of patience. Adaptive quantization and QAT made it happen, LittleBit, Rotation, etc just made it worse. HTML/JS packing with minify, zopfli, and a bunch of structural script changes helped with the result.
  • Was this all made by you? Not at all. This would not have been possible without the 81 KB (FP32) MacroStory model. I "just" tinkered around with it somewhat.

Details for those who're interested:

  • Size-reduction rules are very different for a tiny model than for a large one.
  • The embeddings are 60% of the model size, while the tensors are the largest part in a normal-sized LLM.
  • The tensors of tiny models are extremely sensitive to quantization, especially with the Ouro looping transformer format of this LLM. The embeddings had some room though.
  • It's mathematically impossible to save space with 1 bit quants or less with the LittleBit approach here, as the model matrices are just too small for that - the gains would get eaten up by the overhead of the correction data.
  • GGUF K quants would also not help for the same reason: the model matrices are simply too small for the introduced overhead.
  • Even if some tensor size could be reduced by LittleBit: the 3 KB of quantization gain would be eaten up by 3 KB of newly required safetensors metadata. Nobody thinks about metadata sizes in normal-sized LLMs.
  • Yet aside from that: A modest 4 bit LittleBit quant would still break the model. The currently chosen approach meanwhile comes with a convenient 0.04 KLD. This already causes a very occasional duplicate sentence or non-matching story start.
  • The MacroStory model has a handful of dead tokens that could be exploited for further size reduction, as they'd never make it through the sampler on their own.

Since you've arrived down here, there's a bonus for you: Local Python inference (YMMV) and the non-packed HTML. Just save this image to disk (important: "Download original image"). I originally wanted to use it here or on imgur, but both wouldn't let me. Open with 7-Zip, WinRAR, etc to unpack it. Or on the console - even on Windows - use either:

tar -xf MS256story.png
7z e MS256story.png
python -m zipfile -e MS256story.png

posted Fri, 09 Oct 2026 15:27:09 GMTseen 1 time
open on reddit ↗ 💬 12