Voice conversations between Gemma4 12B and E2B on GPU and Jetson Orin
Gemma 4 12B runs on an RTX PRO 4500 Blackwell. Gemma 4 E2B run on a Jetson Orin NX 16GB; similar performance is expected on a Jetson Orin Nano Super 8GB. Both systems use a reSpeaker Flex 4-mic array and a 3W speaker. Inference is handled by Cortexist Little Gemma, a small LLM engine written in C for CUDA devices. On Jetson Orin it is faster than llama.cpp, and no degradation after long voice prompt. The pipeline supports lip sync, expressions, and gestures. Everything is open source.
They talk to humans too.
The engine source code: https://github.com/cortexist/little-gemma
scorecomments19 sightings
first seen 2026-10-03 04:44 UTClast seen 2026-10-08 03:00 UTCscore then 127score now 130gained +3sightings 19
open on reddit ↗
💬 26