← r/LocalLLaMA
▲
0
 
13👁
r/LocalLLaMA · u/VerityAISolutions · 7d ago

I built an OpenAI-compatible server that runs Gemma 4 E4B on a Pixel 10 Pro XL (Tensor G5) — fully offline, ~11 tok/s decode, Tailscale-encrypted option, Apache 2.0

What it is: an Android app (PixelUnlockGPU) that turns a Pixel into an OpenAI-compatible HTTP server. Standard \/v1/chat/completions\ with streaming, so it talks to TypingMind or any OpenAI client directly — no cloud, no subscription, model runs entirely on-device via LiteRT-LM.

Device: Pixel 10 Pro XL (Tensor G5, 16 GB shared LPDDR). Model: Gemma 4 E4B instruct, GPU bundle, 2.97 GB, SHA-256 verified on download. Context window capped at 32k.

Measured numbers (not benchmarks — real on-device measurements):

\- \~11 tok/s steady-state decode (first-token-to-last over a \~300-word generation)

\- Follow-up turns in \~1.7 s: the server auto-reuses the KV prefix across turns, so stateless clients like TypingMind don't re-prefill history

\- Engine warm build \~12 s once per model change (visible in-app, split out of the metrics on purpose)

\- Short replies read slower than 11 tok/s because warm + prefill dominate the window — the UI separates decode tok/s from prefill ms so nobody has to guess

Security/access: three independent modes — loopback only, raw LAN (unencrypted), or Tailscale (binds the CGNAT tailnet IP; WireGuard end-to-end from e.g. a laptop on the same tailnet; degrades to loopback if the VPN drops). Verified with a real chat completion over the tailnet from a MacBook.

Honest limitations:

\- NPU path aborts on stock Tensor G5 firmware — GPU is the shipping backend (documented with the full investigation)

\- Android blocks named GPU temp sensors for normal apps, so the in-app gauge shows OS thermal headroom instead of °C

\- No real token counts anywhere — LiteRT-LM exposes none, so usage is estimated at \~4 chars/token and labeled as such

\- \stop\ sequences rejected explicitly (engine has no per-request stop API); client-side emulation is a filed issue

Why not llama.cpp/ollama on the phone: wanted the official LiteRT-LM GPU path on Tensor specifically, an always-on Android service (Ktor/Netty), and the OpenAI wire so existing clients just work. Happy to add a GGML backend if people want it — the engine layer is abstracted.

Built on, with credit: server/inference foundation derived from mlnomadpy/localllm (Apache 2.0); Tensor G5 runtime knowledge and the prebuilt dispatch lib from jegly/Box (Apache 2.0). Full attribution in NOTICE + per-file headers. Apache 2.0, contributions welcome — there are labeled good-first-issues (usage block, stop-sequence emulation, docs).

Repo + APKs (v0.1.0/v0.1.1 on releases): https://github.com/cannitellinicholas-spec/PixelUnlockGPU

Happy to answer anything about Tensor G5 quirks — I've done more Gate-2 debugging than I planned to.

1 0 0 10/3 06:28 10/7 09:43 UTC
scorecomments13 sightings
first seen 2026-10-03 06:28 UTClast seen 2026-10-07 09:43 UTCscore then 0score now 0gained 0sightings 13
open on reddit ↗ 💬 2 (+1)