All WordPress HTML Templates Forms & Webhooks AI & Tools
AI & Tools

Browser AI: On-Device Models and What They Unlock

A practitioner’s guide to browser AI in 2026: Chrome’s built-in APIs vs WebGPU with transformers.js, real model sizes, cold start budgets and fallback

A real working moment illustrating the theme of an article about browser ai. Wide 16:9 banner, one strong focal point, magazine editorial quality, authentic and unstaged.

We shipped a documentation search for a client last year that runs entirely in the browser. No search API, no vector database, no per-query cost. A 23 MB quantized embedding model downloads once, gets cached, and answers queries in about 30 ms on a warm session. The whole thing costs us the bandwidth of one medium-sized hero image per new visitor.

That project is the honest case for browser AI: narrow tasks, small models, tangible savings. The dishonest case is the demo where someone runs an 8B parameter chat model on their M3 Max, films it, and implies your users can do the same. They can’t. Half of them are on a Windows laptop with integrated graphics and 8 GB of RAM.

Both stacks are real now. Chrome ships task-specific AI APIs in stable. WebGPU is available in all three major engines. The question stopped being “is this possible” somewhere in 2025 and became “which of these two completely different approaches fits the job”.

Key Takeaways

  • There are two separate stacks: Chrome’s built-in task APIs (Translator, Summarizer, Prompt) backed by Gemini Nano, and bring-your-own-model via WebGPU using transformers.js, WebLLM or ONNX Runtime Web. They have almost nothing in common operationally.
  • Chrome’s built-in model costs you zero bytes of bundle but has hard hardware gates: roughly 22 GB free disk, more than 4 GB of VRAM, and an unmetered connection for the initial multi-gigabyte download. Expect a large share of real traffic to return unavailable.
  • Small purpose-built models win. A 23 MB int8 embedding model doing semantic search or classification is production-viable today; a 2 GB general chat model in a web page usually is not.
  • Always design the no-AI path first. Feature detection, an availability check, a server fallback and a plain UI that works when the model never loads. Treat on-device inference as progressive enhancement, never as a dependency.
  • The genuine unlocks are privacy-shaped: client-side PII redaction before submit, offline translation, local transcription, and search over documents that legally cannot leave the device.

Browser AI in 2026 means two different stacks

The first stack is the built-in one. Chrome bundles Gemini Nano and exposes it through task-scoped JavaScript APIs: Translator and LanguageDetector and Summarizer reached stable in Chrome 138, with Writer, Rewriter, Proofreader and the general-purpose LanguageModel (the Prompt API) moving through origin trials and extension availability since. You ship no model weights. The browser manages the download, the cache, the updates and the memory.

The second stack is bring-your-own. You pick weights from Hugging Face, quantize them, and run them through WebGPU with transformers.js v3, WebLLM or ONNX Runtime Web. You control the model, the version, the quantization and the behaviour. You also own every megabyte of download and every GPU out-of-memory crash.

Choosing between them is not a preference. If your task is translation, summarization or proofreading and you only need it on Chromium, the built-in API is the obvious choice: free, zero bundle weight, decent quality. If you need cross-browser support, a specific model, embeddings, audio, images, or reproducible output across sessions, you’re in the second stack whether you like it or not.

What the built-in APIs actually do well

They do one thing exceptionally well: they let you add a feature that costs nothing when it’s unavailable. The availability check is the entire API surface that matters.

// Every built-in API follows the same shape.
// availability() returns: 'unavailable' | 'downloadable' | 'downloading' | 'available'
if (!('Summarizer' in self)) return null;

const status = await Summarizer.availability({
  type: 'key-points',
  format: 'markdown',
  outputLanguage: 'en'
});

if (status === 'unavailable') return null;   // hardware gate failed, bail silently

const summarizer = await Summarizer.create({
  type: 'key-points',
  length: 'short',
  // Only fires when status was 'downloadable'. Multi-GB on first run.
  monitor(m) {
    m.addEventListener('downloadprogress', e => {
      console.log(${Math.round(e.loaded * 100)}%);
    });
  }
});

const summary = await summarizer.summarize(articleText, {
  context: 'For a technical audience already familiar with the product.'
});
summarizer.destroy(); // release the session; do not leak these

Two things bite people here. First, create() can trigger a download measured in gigabytes. Chrome requires roughly 22 GB of free disk space plus more than 4 GB of VRAM before it will even consider it. Never call create() on page load. Gate it behind a user gesture.

Second, the Prompt API is not a frontier model and will not behave like one. Structured output is often the difference between shipping and not shipping. It’s supported through responseConstraint:

const session = await LanguageModel.create({
  temperature: 0.1,   // near-deterministic; defaults are too creative for extraction
  topK: 1,
  initialPrompts: [{
    role: 'system',
    content: 'You classify support messages. Reply only with the schema.'
  }]
});

const schema = {
  type: 'object',
  required: ['category', 'urgency'],
  additionalProperties: false,
  properties: {
    category: { type: 'string', enum: ['billing', 'bug', 'feature', 'other'] },
    urgency: { type: 'string', enum: ['low', 'normal', 'high'] }
  }
};

const raw = await session.prompt(userMessage, { responseConstraint: schema });
const result = JSON.parse(raw);

// Sessions have a finite token budget. Check before reusing.
console.log(session.inputUsage, '/', session.inputQuota);
session.destroy();

Nano’s context window is small, in the low thousands of tokens, so long-document work needs chunking. Push a 40 KB article into a session and you’ll hit the quota and get a rejected promise, not a graceful truncation. The prompt patterns that survive model updates apply here even harder than server-side, because you cannot pin the version of the model Chrome ships.

Bring your own model: WebGPU AI in practice

WebGPU AI stopped being a Chrome-only story. Chrome has shipped it since 113, Firefox enabled it on Windows in 141, and Safari 26 brought it to macOS and iOS. Coverage is genuinely broad now, though the shader-f16 feature (which most fast quantized kernels want) is not universal, so check for it explicitly.

async function gpuProfile() {
  if (!navigator.gpu) return { ok: false, reason: 'no-webgpu' };
  const adapter = await navigator.gpu.requestAdapter();
  if (!adapter) return { ok: false, reason: 'no-adapter' };
  return {
    ok: true,
    f16: adapter.features.has('shader-f16'),
    // Single-buffer ceiling. On low-end mobile this can be ~128 MB,
    // which rules out loading big weight tensors in one allocation.
    maxBuffer: adapter.limits.maxBufferSize
  };
}

For anything that isn’t text generation, transformers.js is the pragmatic choice. Embeddings, zero-shot classification, token classification for PII, Whisper for transcription, background removal. These models are tens of megabytes, not gigabytes.

import { pipeline, env } from '@huggingface/transformers';

env.allowLocalModels = false; // fetch from CDN, cache in Cache Storage

const embed = await pipeline(
  'feature-extraction',
  'Xenova/all-MiniLM-L6-v2',
  { device: 'webgpu', dtype: 'q8' } // ~23 MB int8, 384-dim output
);

const out = await embed(['reset my password'], { pooling: 'mean', normalize: true });
const vector = out.tolist()[0]; // cosine similarity against prebuilt doc vectors

Run this in a Web Worker. Not optional. Tokenization and post-processing are CPU-bound JavaScript. On the main thread they will blow your Interaction to Next Paint apart even though the matrix multiplies happen on the GPU.

For actual chat, WebLLM is the mature option and its prebuilt model list is a useful reality check on sizes. A 0.5B instruct model at 4-bit lands around 400 to 500 MB. An 8B model at 4-bit is roughly 4.5 GB. On an M-series Mac the 8B decodes at a comfortable reading pace. On a mid-range Windows laptop with integrated graphics, the same model is a slideshow, assuming it loads at all. That gap is the whole story of on-device generation.

The cold start is the product decision

Everyone benchmarks tokens per second. Almost nobody budgets the first load, which is what users actually experience.

Weights land in Cache Storage, and Cache Storage is evictable. On iOS in particular, storage gets reclaimed aggressively, so a returning visitor may pay the download twice. Request persistence and check what you’ve actually been granted:

const persisted = await navigator.storage.persist();      // may be false without install/engagement
const { quota, usage } = await navigator.storage.estimate();
// Refuse to start a 2 GB download when quota is 300 MB.
if (quota - usage < MODEL_BYTES * 1.4) useServerFallback();

Our rule on client work is blunt. Under 50 MB: download lazily on first use, say nothing. From 50 MB to about 300 MB: ask, show a progress bar, remember the choice. Above 300 MB: it has to be an explicit opt-in with a visible benefit the user requested, typically “work offline” or “keep this data on my device”. Silently pulling half a gigabyte because someone hovered over a textarea is how you end up in a support thread about a mobile data bill.

What an on device LLM is good at, and where it falls over

An on device LLM handles the boring middle of the stack well. Classification and routing. Extracting structured fields from a messy paste. Rewriting a sentence to a fixed tone. Detecting language. Semantic similarity. Redacting an email address or a card number before a form hits the network. These tasks share short inputs, short outputs, and tolerance for a small model.

It falls over on:

  • Anything needing world knowledge. A 0.5B model does not know your pricing, your API, or last quarter’s changelog. Retrieval on-device plus generation on-device compounds two weak links.
  • Long context. Small windows and small quotas mean chunking, and chunking means orchestration logic that is often more code than the feature is worth.
  • Consistency across users. Different browsers, different model versions, different quantizations. The same prompt genuinely produces different output for different visitors, which makes support and QA harder than people expect.
  • Battery and thermals. Sustained GPU inference on a phone throttles within a minute or two and drains noticeably. Fine for a one-shot summarize. Bad for a chat UI someone keeps open.

If the output is customer-facing, the same guardrails you’d apply to a server-side model still apply. The difference is you cannot inspect the logs, because there aren’t any. Constrain output with a schema, validate it, and render a static fallback when validation fails.

Hybrid routing is the pattern that holds up

The architecture that has survived contact with real client sites is a three-tier cascade, evaluated per feature and not per app:

  1. Tier 0, no model. The feature works in a degraded but genuinely useful form. Keyword search instead of semantic search. A regex for obvious PII. This tier ships first and always ships.
  2. Tier 1, on-device. If the built-in API reports available, or the GPU profile passes and the small model is cached, upgrade in place. No layout shift, no spinner.
  3. Tier 2, server. For heavy or quality-critical work, or when the user explicitly asks for the better answer. Cheap to run because tiers 0 and 1 already absorbed most of the volume.

One concrete example worth copying: client-side redaction before submit. Run a small token-classification model over the message body, mask what it flags, and show the user exactly what will be sent. The raw text never leaves the device, which changes the GDPR conversation entirely. It pairs cleanly with a form endpoint like WebForms on a static site, since the sensitive processing happens before the POST and there is no backend to audit.

Measure four numbers, not one

Tokens per second is the least useful metric on the list. Track these instead, in real user monitoring and not just on your own machine:

  • Availability rate. What percentage of sessions reach available or pass your GPU profile? Below 40 percent, the built-in path is a bonus feature and your server path is the product.
  • Time to first token, cold and warm. Cold includes download and GPU pipeline compilation. Warm is the only number that should be sub-second.
  • Cache hit rate on return visits. If this is low, your storage assumptions are wrong and users are re-downloading.
  • Long tasks introduced. Run the feature in a Performance trace. Any task over 200 ms on the main thread means work escaped the worker.

Set explicit abort behaviour too. Every built-in API accepts an AbortSignal, and every worker-based pipeline should. If inference exceeds your budget, cancel it and fall back rather than leaving a spinner running while a phone heats up in someone’s hand. The same discipline applies as with prompts you expect to outlive a model update: constrain, validate, and have a defined behaviour for failure.

Frequently Asked Questions

Is browser AI supported outside Chrome?

WebGPU is, so the bring-your-own-model stack runs in Chrome, Edge, Firefox on Windows and Safari 26 on macOS and iOS. The built-in task APIs (Translator, Summarizer, Prompt) are Chromium-only, since they depend on Gemini Nano shipping with the browser. Treat the built-in APIs as a Chromium enhancement and WebGPU as the portable baseline.

How large a model can I realistically ship to a web page?

Under 50 MB for a feature you enable by default, and up to a few hundred megabytes for an explicit opt-in. Multi-gigabyte chat models only make sense in an installed PWA or a Chrome extension where the user has already accepted an install step. The failure mode is not slow inference, it’s the user closing the tab during the download.

Does on-device inference remove my GDPR obligations?

It removes the processing obligations for data that genuinely never leaves the device, which is a real and significant reduction. It does not remove obligations for anything you subsequently transmit, store, or log. Document which tier of your cascade handled the data, because “it ran locally” is only defensible if you can show the network requests that were not made.

WebGPU or WebNN for AI in the browser?

WebGPU today. WebNN is the more interesting long-term target because it can reach NPUs and dedicated accelerators rather than just the GPU, but it remains behind flags with no default-on browser as of early 2026. Build on WebGPU and keep your inference behind an interface so swapping the backend later is a one-file change.

Can I get deterministic output from an on-device model?

Not across users. Setting temperature near zero and topK to 1 gets you close within a single browser and model version, but different quantizations and hardware paths produce different floating-point results. If you need identical output for identical input, run it server-side with a pinned model and cache the result.

Here’s the decision. Pick one narrow task where the data is sensitive or the volume is high enough that per-call API costs annoy you, build the no-model version first, then add a cached small model behind an availability check and measure how many of your users actually reach it. That number will tell you honestly whether browser inference is a feature or a science project on your particular traffic. Everything else follows from it.

browser ai chrome prompt api availability client side pii redaction javascript gemini nano browser requirements on device llm transformers.js webgpu embeddings webgpu ai webllm model size browser