← r/LocalLLaMA
▲
0
 
12👁
r/LocalLLaMA · u/cortexist · 11d ago

A hybrid model of Gemma4 with a JEV-like decision head in multi-speaker voice conversation

post image

The human brain is neither an LLM nor a JEV. In a crowded market you hear a lot of speech and answer almost none of it. The ongoing question is not “what should I say?” It is “you talking to me?" and "should I say anything at all?”

This live voice demo showcases three hardware tiers—the Blackwell 4500, Jetson Orin NX 16GB, and Jetson Orin Nano 8GB—solving this exact problem. By splitting the workload between a lightweight decision head for turn-taking and a Gemma 4 pipeline for text generation, the setup delivers highly responsive, low-latency vocal interaction.

EDIT: repo (the latest code yet published) https://github.com/cortexist/little-gemma

1 0 0 10/3 06:30 10/7 08:01 UTC
scorecomments12 sightings
first seen 2026-10-03 06:30 UTClast seen 2026-10-07 08:01 UTCscore then 0score now 0gained 0sightings 12
open on reddit ↗ 💬 20