Image-text retrieval with EmbeddingGemma 2's vision tower, running in the browser on WebGPU
EmbeddingGemma 2 came out this week. It maps images and text into one 768-dim space, so you can search photos by describing them. I ported its text and vision towers to ruNNtime, a WebGPU inference library in TypeScript, and made a small photo gallery where search runs entirely on your GPU in the browser.
ruNNtime also supports plenty of other vision-like models, and you can play with them in the interactive docs
scorecomments19 sightings
first seen 2026-10-07 11:19 UTClast seen 2026-10-09 07:48 UTCscore then 27score now 77gained +50sightings 19
open on reddit ↗
💬 7 (+4)