← r/LocalLLaMA
▲
77
+50
19👁
r/LocalLLaMA · u/FinancialAd1961 · 2d ago

Image-text retrieval with EmbeddingGemma 2's vision tower, running in the browser on WebGPU

post image

EmbeddingGemma 2 came out this week. It maps images and text into one 768-dim space, so you can search photos by describing them. I ported its text and vision towers to ruNNtime, a WebGPU inference library in TypeScript, and made a small photo gallery where search runs entirely on your GPU in the browser.

ruNNtime also supports plenty of other vision-like models, and you can play with them in the interactive docs

source: https://github.com/software-mansion/runntime

77 0 77 10/7 11:19 10/9 07:48 UTC
scorecomments19 sightings
first seen 2026-10-07 11:19 UTClast seen 2026-10-09 07:48 UTCscore then 27score now 77gained +50sightings 19
open on reddit ↗ 💬 7 (+4)