Run Google Gemma LLMs directly in your browser — no server, no backend, no data leaving your device.
Browser LLM is a React web application that runs large language models (LLMs) entirely inside your browser tab using Transformers.js and WebGPU. There is no server, no API calls, and no data ever leaves your machine. After the first download the app works completely offline.
Modern browsers expose a GPU compute API called WebGPU. Transformers.js compiles models to the ONNX format and runs them through ONNX Runtime Web, which translates operations into WebGPU shader calls. The result is near-native GPU inference speed inside a sandboxed browser tab.
Your browser tab
+-- React UI
+-- Transformers.js (v4)
+-- ONNX Runtime Web
+-- WebGPU --? GPU (DirectX12 / Metal / Vulkan)
Models are downloaded once from Hugging Face and cached permanently in the browser's Cache Storage (survives browser restart, works offline).
| Model | Type | Size | Use |
|---|---|---|---|
onnx-community/gemma-3-1b-it-ONNX |
Text only | ~500 MB | Fast chat, Q&A |
onnx-community/gemma-3n-E2B-it-ONNX |
Text + Image | ~3.5 GB | Multimodal understanding |
Both models are gated — you need a free Hugging Face account and must accept the Gemma license before they can be downloaded.
{
vision_encoder: 'fp32', // ~1.2 GB — no smaller variant available
audio_encoder: 'q4', // ~451 MB
embed_tokens: 'q4', // ~1.6 GB
decoder_model_merged: 'q4f16' // ~1.44 GB — fastest on WebGPU
}- Zero backend — model runs entirely in the browser via WebGPU
- Two models — switch between text-only (fast) and multimodal (text + image)
- Image upload — drag or select an image when using Gemma 3n E2B
- Persistent cache — models cached after first download, works offline
- Model switcher — switch models without losing cached instances
- HF token UI — enter your Hugging Face token in the UI (stored in localStorage only)
- WebGPU detection — auto-detects GPU support, falls back to WASM
- Privacy-first AI — all inference local, nothing sent to any server
- Offline environments — works without internet after first model download
- Enterprise / air-gapped — no external API dependency or data exposure
- Education — demonstrate how LLMs work without infrastructure
- Prototyping — test prompts and model behaviour without API costs
- Image analysis — describe, caption or question images locally (Gemma 3n E2B)
| Requirement | Detail |
|---|---|
| Browser | Chrome 113+ or Edge 113+ |
| WebGPU | Required for GPU acceleration (WASM fallback is very slow) |
| RAM | 4 GB minimum, 8 GB+ recommended for E2B |
| VRAM | 2 GB+ (shared or dedicated) |
| Storage | ~500 MB (1B model) or ~3.5 GB (E2B model) browser cache |
- Node.js 18+
- npm 9+
- Chrome 113+ or Edge 113+ (WebGPU enabled)
- A free Hugging Face account with Gemma license accepted
# 1. Clone the repo
git clone https://github.com/your-username/WebBrowserLLM.git
cd WebBrowserLLM
# 2. Install dependencies
npm install
# 3. Start the dev server
npm run dev
# 4. Open in Chrome
# Navigate to http://localhost:5173- Create a free account at huggingface.co
- Accept the Gemma license at:
- Create a Read token at huggingface.co/settings/tokens
- Paste the token into the app UI when prompted
The token is stored only in your browser's localStorage. It is never sent anywhere except Hugging Face to authenticate the model download.
WebBrowserLLM/
+-- index.html # Entry HTML (includes Tailwind CDN)
+-- vite.config.js # Vite config (react-swc plugin)
+-- package.json
+-- src/
+-- App.jsx # Root: shows ModelLoader or ChatInterface
+-- main.jsx # React entry point
+-- index.css # Base styles
+-- components/
¦ +-- ModelLoader.jsx # Model selector + download UI
¦ +-- ChatInterface.jsx # Main chat layout + model switcher
¦ +-- InputBox.jsx # Text input + image upload
¦ +-- MessageList.jsx # Message rendering with image thumbnails
¦ +-- DeviceInfo.jsx # Device/GPU capability display
¦ +-- WebGPUWarning.jsx # Warning when WebGPU unavailable
+-- hooks/
¦ +-- useTransformersModel.js # Model loading, generation, caching logic
+-- store/
¦ +-- appStore.js # Zustand global state
+-- utils/
+-- modelConfig.js # Model definitions, WebGPU detection
| Library | Version | Purpose |
|---|---|---|
| React | 18 | UI framework |
| Vite | 4 | Build tool (react-swc, no Babel) |
| @huggingface/transformers | v4 (next) | LLM inference via ONNX Runtime Web |
| Zustand | 4 | Global state management |
| Tailwind CSS | 3 (CDN) | Styling |
| WebGPU | Browser API | GPU acceleration |
npm run build
# Output in dist/ — serve with any static host (Nginx, GitHub Pages, Netlify, etc.)Note: Users must access the app over HTTPS (or localhost) for WebGPU to be available.
| Setup | Speed |
|---|---|
| WebGPU (GPU) | ~5–15 tokens/sec |
| WASM (CPU fallback) | ~1–3 tokens/sec |
| Ollama / native | 30–100 tokens/sec |
Browser inference is ~5–10x slower than native runtimes (Ollama, llama.cpp) due to WebGPU abstraction overhead and ONNX kernel quality. The tradeoff is zero installation and complete privacy.
MIT