← r/LocalLLaMA
▲
9
-1
24👁
r/LocalLLaMA · u/stevyhacker · 9d ago

Five local models, 6.8 GB of weights: my open-source Mac meeting notetaker

https://preview.redd.it/xs01ds930nsh1.png?width=4800&format=png&auto=…

Back in July I shared LokalBot here. It's a free, open-source Mac app that records your meetings and keeps a daily summary of your activity, all on-device.

0.9.2 came out today. Since July I've benchmarked every model in it and swapped most of the defaults for smaller ones. The whole stack is now 6.8 GB:

  • Qwen3-ASR 1.7B (MLX, 8-bit): transcription
  • Nemotron 3 (Core ML): who spoke when
  • Qwen3.5 4B Q4\_K\_M (llama.cpp): notes and action items
  • Harrier 0.6B Q8\_0: search embeddings
  • LFM2.5 1.2B Q4\_K\_M: autocomplete in any app
  • Apple Vision: screen OCR (opt-in)

A few numbers from my M4 Max (48 GB):

  • 26-min meeting to finished notes in 33 s warm, \~85 tok/s decode
  • Speaker error went from 43.4% to 14.6% DER on AMI. That's against my old pyannote setup, so it says more about my config than about pyannote.
  • Autocomplete p95 went from 1.83 s (Gemma 4 E4B) to 0.49 s

There's also a read-only MCP server and CLI, off by default, so Claude Code or any other MCP client can pull context from your meetings.

I don't have any 16 GB or other M series numbers yet. If you've got one of those, especially M5 or M6 I'd love to see what you get.

I also tried MiniCPM5 2B for notes. It was smaller and faster, but it got stuck repeating itself on one summary and assigned action items to the wrong person. I kept Qwen3.5 4B as the default as saving a few seconds wasn’t worth getting who agreed to do what wrong.

11 0 9 10/3 06:29 10/7 09:51 UTC
scorecomments24 sightings
first seen 2026-10-03 06:29 UTClast seen 2026-10-07 09:51 UTCscore then 10score now 9gained -1sightings 24
open on reddit ↗ 💬 4 (+1)