TTS silently dropped 17% of a passage and nobody could hear it — so I built a local audiobook tool that transcribes every line back
I've been building VoxStage, a local script-to-voice workstation for Apple Silicon Macs. Paste a chapter of prose with no speaker labels, and it gives you a multi-voice reading you can audition line by line, fix, redo and export. Everything runs on the Mac: no account, no cloud API, no telemetry.
Why it exists: in an earlier local voice-cloning test, a long passage came out fluent and natural — and 40 characters (about 17%) from the middle were simply gone. The remaining text still read as a normal sentence, so nobody could hear it. That changed two design rules:
- Generate sentence by sentence, never a whole passage at once.
- Transcribe every generated line back with a local recogniser (whisper.cpp) and diff it against the script. Disagreements are flagged for your ear, never auto-corrected.
The stack:
\- Speech: Qwen3-TTS on MLX (0.6B / 1.7B preset voices, voice design from a description, cloning from a recording you confirm you have the rights to)
\- Who says what: a local LLM via llama.cpp (Qwen3-14B, or Qwen3-30B-A3B on 32 GB) drafts the speaker for each line; program-side rules on top; you review
\- Read-back check: whisper.cpp
Measured on my M2 Max 32 GB, Pride and Prejudice ch. 1: speaker draft for 35 units in 14.8 s; 28 lines → 144.7 s of audio synthesised in 51 s (RTF 0.358, preset-voice path); read-back check 28 s.
Honest limits:
\- The speaker draft is a draft. In my evaluation most scenes needed at least one correction, so the review step is the product, not a formality.
\- Chinese and English only for now.
\- Install is still developer-style (Homebrew + terminal, \~30 min mostly model downloads) and only verified on my own Mac. A signed one-click installer is in progress.
Other things it does: editing one sentence regenerates only that sentence; subtitles (SRT/VTT) timed from the actual audio; an FCP7 XML timeline that imports into DaVinci Resolve; long texts kept as a book with chapters inheriting the cast.
Samples (longer ones first): https://houjun.dev/voxstage/#listen
Code (AGPL-3.0): https://github.com/hera2019/VoxStage
I'd especially like to hear:
\- Which local models you've found best at speaker attribution in fiction
\- Whether anyone has seen the same silent-skip behaviour with other TTS models
\- If you try the install on a Mac other than an M2 Max, whether it works