
WebGPU and Client-Side AI: In-Browser Inference with JavaScript and TypeScript
How to run machine learning models directly in the browser with WebGPU and Transformers.js: model loading, Web Workers, WASM fallback and memory limits.
Until recently, “adding AI to a web app” meant exactly one thing: an HTTP call to an API, with network latency, per-token cost, and user data leaving the device. For many use cases that’s still the right call. But for a growing share of features — semantic search over local content, text classification, OCR, background removal, transcription — the model can run directly in the browser, on the user’s GPU, through WebGPU.
The result: zero network latency after the first load, zero server cost for inference, and data that never leaves the device. In exchange, you’re the one managing model downloads, memory, and compatibility across wildly different devices. Let’s see how, with real TypeScript code.
Why WebGPU changes things
WebGL already let you use the GPU from the browser, but it was an API designed for drawing triangles: doing general-purpose compute meant disguising your data as textures. WebGPU instead exposes compute shaders as first-class citizens, with memory buffers, compute pipelines, and a model much closer to Vulkan/Metal/Direct3D 12. That’s exactly the primitive the matrix multiplications at the core of every neural network need.
In practice you’ll almost never write shaders by hand: runtimes like ONNX Runtime Web and libraries like Transformers.js (which uses it under the hood) translate the model graph into WebGPU operations for you. When WebGPU isn’t available, they fall back to WebAssembly on the CPU — slower, but it works everywhere.
ℹ️ Support in 2026
WebGPU has been available in Chrome and Edge since 2023, in Safari since Safari 26, and in Firefox (first on Windows, then on other platforms). Coverage is broad but not universal — especially on Linux and older mobile GPUs — so feature detection with a fallback is mandatory, not optional.
The stack: what to pick
| Library | Level | When to use it |
|---|---|---|
@huggingface/transformers |
High (ready-made pipelines) | Embeddings, classification, translation, OCR, speech-to-text with Hugging Face models |
onnxruntime-web |
Mid (ONNX sessions) | You have your own model exported to ONNX and want control over input/output tensors |
@mlc-ai/web-llm |
High (OpenAI-style API) | Conversational LLMs (Llama, Qwen, Phi…) entirely in the browser |
For this article’s example we’ll use Transformers.js: it’s the fastest way in, and it covers the use case most apps actually benefit from — semantic search.
The use case: local semantic search
Goal: the user types “how do I cancel a fetch request” and finds the document titled “AbortController and cancellation signals”, even with no words in common. We need an embedding model that turns each text into a numeric vector; texts with similar meaning produce vectors that are close together.
We’ll use Xenova/all-MiniLM-L6-v2: small (about 23 MB quantized to 8 bits), fast, 384-dimensional vectors. For non-English content, a multilingual model like Xenova/paraphrase-multilingual-MiniLM-L12-v2 is a better fit, at a somewhat larger size.
npm install @huggingface/transformers
npm install -D @webgpu/types
@webgpu/types adds the type definitions for navigator.gpu, which TypeScript doesn’t yet ship in its default DOM lib:
// tsconfig.json
{
"compilerOptions": {
"types": ["@webgpu/types"]
}
}
Rule number one: inference goes in a Web Worker
Loading a model means downloading tens of MB, decompressing them, compiling shaders. Even with the GPU doing the heavy lifting, tensor preparation and tokenization run in JavaScript. Doing it on the main thread means freezing the UI for hundreds of milliseconds — exactly the problem I described in the article on Web Workers, and one that shows up as a bad INP in your Core Web Vitals.
Good news: WebGPU is also available inside Dedicated Workers (navigator.gpu exists on WorkerNavigator), so everything — download, compilation, inference — can live off the main thread.
Let’s start with the message protocol, typed as a discriminated union so the main thread and the worker can’t drift apart:
// embed.protocol.ts
export type WorkerRequest =
| { type: "load" }
| { type: "embed"; id: number; texts: string[] };
export type WorkerResponse =
| { type: "progress"; file: string; progress: number }
| { type: "ready"; device: "webgpu" | "wasm" }
| { type: "result"; id: number; vectors: Float32Array; dims: number }
| { type: "error"; id?: number; message: string };
The worker
// embed.worker.ts
import { pipeline, type FeatureExtractionPipeline } from "@huggingface/transformers";
import type { WorkerRequest, WorkerResponse } from "./embed.protocol";
const MODEL_ID = "Xenova/all-MiniLM-L6-v2";
let extractor: Promise<FeatureExtractionPipeline> | null = null;
let device: "webgpu" | "wasm" = "wasm";
const post = (msg: WorkerResponse, transfer: Transferable[] = []) =>
self.postMessage(msg, { transfer });
async function detectDevice(): Promise<"webgpu" | "wasm"> {
if (!("gpu" in navigator)) return "wasm";
const adapter = await navigator.gpu.requestAdapter();
return adapter ? "webgpu" : "wasm";
}
function getExtractor() {
// Singleton: one model instance, even if multiple "load" messages arrive
extractor ??= (async () => {
device = await detectDevice();
const pipe = await pipeline("feature-extraction", MODEL_ID, {
device,
// q8 is much faster on WASM; on the GPU fp32 avoids precision issues
dtype: device === "webgpu" ? "fp32" : "q8",
progress_callback: (p) => {
if (p.status === "progress") {
post({ type: "progress", file: p.file, progress: p.progress });
}
},
});
return pipe;
})();
return extractor;
}
self.onmessage = async (event: MessageEvent<WorkerRequest>) => {
const msg = event.data;
try {
if (msg.type === "load") {
await getExtractor();
post({ type: "ready", device });
return;
}
if (msg.type === "embed") {
const pipe = await getExtractor();
const output = await pipe(msg.texts, { pooling: "mean", normalize: true });
const vectors = output.data as Float32Array;
const dims = output.dims[1];
// Transfer the buffer instead of copying it: zero-copy between threads
post({ type: "result", id: msg.id, vectors, dims }, [vectors.buffer]);
}
} catch (err) {
post({
type: "error",
id: msg.type === "embed" ? msg.id : undefined,
message: err instanceof Error ? err.message : String(err),
});
}
};
A few non-obvious choices:
pooling: "mean", normalize: trueproduces a single vector per text, already normalized to length 1. That makes cosine similarity equal to a plain dot product — cheaper to compute.- The buffer is transferred, not copied. With hundreds of documents at 384 floats each, skipping the structured clone matters. After
postMessage,vectorsin the worker is unusable (detached): that’s intended. - The
dtypedepends on the device. 8-bit quantization cuts size and time on the CPU; on the GPU the speed difference is smaller andfp32is the most robust choice. If the adapter supports theshader-f16feature,fp16is a good middle ground.
The main-thread client
// embedder.ts
import type { WorkerRequest, WorkerResponse } from "./embed.protocol";
export class Embedder {
private worker = new Worker(new URL("./embed.worker.ts", import.meta.url), {
type: "module",
});
private nextId = 0;
private pending = new Map<
number,
{ resolve: (v: Float32Array[]) => void; reject: (e: Error) => void }
>();
readonly ready: Promise<"webgpu" | "wasm">;
constructor(onProgress?: (file: string, progress: number) => void) {
this.ready = new Promise((resolve, reject) => {
this.worker.onmessage = (event: MessageEvent<WorkerResponse>) => {
const msg = event.data;
switch (msg.type) {
case "progress":
onProgress?.(msg.file, msg.progress);
break;
case "ready":
resolve(msg.device);
break;
case "result": {
const out: Float32Array[] = [];
for (let i = 0; i < msg.vectors.length; i += msg.dims) {
out.push(msg.vectors.subarray(i, i + msg.dims));
}
this.pending.get(msg.id)?.resolve(out);
this.pending.delete(msg.id);
break;
}
case "error": {
const error = new Error(msg.message);
if (msg.id === undefined) {
reject(error);
} else {
this.pending.get(msg.id)?.reject(error);
this.pending.delete(msg.id);
}
}
}
};
});
this.send({ type: "load" });
}
private send(msg: WorkerRequest) {
this.worker.postMessage(msg);
}
async embed(texts: string[]): Promise<Float32Array[]> {
await this.ready;
const id = this.nextId++;
return new Promise((resolve, reject) => {
this.pending.set(id, { resolve, reject });
this.send({ type: "embed", id, texts });
});
}
dispose() {
this.worker.terminate();
}
}
new URL("./embed.worker.ts", import.meta.url) is the pattern Vite, webpack 5, and other modern bundlers recognize to emit a separate chunk for the worker — no extra configuration.
The actual search
// search.ts
import { Embedder } from "./embedder";
const dot = (a: Float32Array, b: Float32Array) => {
let sum = 0;
for (let i = 0; i < a.length; i++) sum += a[i] * b[i];
return sum;
};
const docs = [
"AbortController and cancellation signals",
"Debounce and throttle in JavaScript",
"Content Security Policy: nonces and strict-dynamic",
"Web Workers: moving work off the main thread",
];
const embedder = new Embedder((file, p) => console.log(`${file}: ${p.toFixed(0)}%`));
console.log("Device:", await embedder.ready);
// Computed once, then optionally persisted to IndexedDB
const docVectors = await embedder.embed(docs);
export async function search(query: string, k = 3) {
const [q] = await embedder.embed([query]);
return docs
.map((text, i) => ({ text, score: dot(q, docVectors[i]) }))
.sort((a, b) => b.score - a.score)
.slice(0, k);
}
console.table(await search("how do I cancel a fetch request"));
For a few thousand documents, a linear scan is more than enough (a few million multiplications, under a millisecond). Beyond that, reach for an approximate index (HNSW) or move the scoring into the worker too.
💡 Tip
Document embeddings don’t change until the text does: store them in IndexedDB along with a content hash and recompute only the ones that changed. The model itself is already cached by Transformers.js through the browser’s Cache API, so from the second load the download disappears.
Handling memory and very different devices
This is where a demo and a production feature part ways. The same code will run on a workstation with 24 GB of VRAM and on a mid-range phone with shared memory. A few checks to run before downloading the model:
// capabilities.ts
export async function inspectDevice() {
const gpu = "gpu" in navigator ? await navigator.gpu.requestAdapter() : null;
const storage = await navigator.storage?.estimate?.();
return {
webgpu: !!gpu,
// Single-buffer limit: a layer's weights must fit inside it
maxBufferSize: gpu?.limits.maxBufferSize ?? 0,
maxStorageBinding: gpu?.limits.maxStorageBufferBindingSize ?? 0,
fp16: gpu?.features.has("shader-f16") ?? false,
// Chromium only, rounded and capped at 8 GB for privacy: a hint, not a number
deviceMemoryGB: (navigator as { deviceMemory?: number }).deviceMemory,
// Space available for caching the model
quotaFreeMB: storage ? ((storage.quota ?? 0) - (storage.usage ?? 0)) / 1e6 : undefined,
};
}
With this information you can apply a tiered strategy:
- WebGPU + plenty of memory → larger model or
fp32/fp16precision. - WebGPU with low limits, or
deviceMemory≤ 4 → small quantized model (q8/q4). - No WebGPU → WASM with a quantized model, and maybe a UI hint that the operation will be slower.
- Not enough storage quota, or a metered connection (
navigator.connection?.saveData) → don’t download anything automatically: ask the user, or fall back to a server API.
Another event to handle is GPU device loss: the driver can reset, the OS can reclaim memory, a background tab can get deprioritized. If you work with WebGPU directly, device.lost is a Promise that resolves when it happens; with Transformers.js or ONNX Runtime you’ll see it as an error during inference. Either way the strategy is the same: terminate the worker, spawn a new one, and reload the model (from cache, so no new download).
⚠ Warning
ONNX Runtime’s multi-threaded WASM fallback uses SharedArrayBuffer, which requires a
cross-origin isolated page (Cross-Origin-Opener-Policy: same-origin and
Cross-Origin-Embedder-Policy: require-corp headers). Without them the runtime still works, but
single-threaded — much slower. Be careful: COEP can break third-party embeds and scripts that don’t
send the right CORP/CORS headers.
What about LLMs?
The same approach extends to generative models with WebLLM, which exposes an OpenAI-compatible API:
import { CreateMLCEngine } from "@mlc-ai/web-llm";
const engine = await CreateMLCEngine("Qwen2.5-1.5B-Instruct-q4f16_1-MLC", {
initProgressCallback: (p) => console.log(p.text),
});
const reply = await engine.chat.completions.create({
messages: [{ role: "user", content: "Summarize what WebGPU is in one sentence." }],
});
console.log(reply.choices[0].message.content);
WebLLM also offers CreateWebWorkerMLCEngine to run everything in a worker behind the same API. But the proportions change: a 1-3 billion parameter model quantized to 4 bits is one or more GB to download and needs about as much GPU memory. It makes sense for tools people use repeatedly (an editor, an internal app), much less for a landing page visited once.
When it makes sense (and when it doesn’t)
Client-side inference pays off when:
- the model is small (tens to a few hundred MB) and reused several times per session;
- the data is sensitive and shouldn’t leave the device (documents, photos, private text);
- you want offline or local-first features;
- request volume would make a pay-per-use API expensive.
The server is still the better choice when you need a large, top-quality model, when target devices are old or underpowered, or when the first use is also the only one and a download of tens of MB isn’t justified.
Frequently asked questions
❓ Does WebGPU work inside a Web Worker?
Yes. navigator.gpu is also exposed on WorkerNavigator in Dedicated Workers, so download,
shader compilation, and inference can all stay off the main thread.
❓ What happens if the browser doesn't support WebGPU?
Transformers.js and ONNX Runtime Web can use WebAssembly on the CPU. Pick device: "wasm" (or let
the runtime fall back) and prefer a quantized model: slower, but it works everywhere.
❓ Is the model downloaded again on every visit?
No. Transformers.js stores model files in the browser’s Cache API, so from the second load they’re read locally. The browser can still evict that cache when storage runs low, so your code must always handle a fresh download.
❓ How big is a browser-friendly model?
For embeddings and classification, 20-100 MB is enough. OCR and speech-to-text models range from a few tens to a few hundred MB. LLMs, even small ones quantized to 4 bits, start at around a GB.
Conclusion
WebGPU has turned in-browser inference from an experiment into a real architectural option. The model code is the easy part — a Transformers.js pipeline is three lines. The real work is everything around it: keeping the main thread free with a Web Worker, choosing model and precision based on the device, handling cache, fallbacks, and device loss. Get that right and you ship AI features with zero latency, zero server cost, and privacy by design.
If performance is what you care about, the starting point is the article on Web Workers in JavaScript; to cleanly cancel searches still in flight while the user keeps typing, the same pattern described in AbortController applies.

Co-Founder & CTO at PAPION. Senior full-stack engineer specializing in React, TypeScript, Node.js, and application security.