Diego Betto
Photo by Jackson Sophat on Unsplash

Diego Betto · September 29, 2026 · 12 min di lettura

WebGPU and Client-Side AI: In-Browser Inference with JavaScript and TypeScript

How to run machine learning models directly in the browser with WebGPU and Transformers.js: model loading, Web Workers, WASM fallback and memory limits.

Condividi:XLinkedInFacebookWhatsApp

Until recently, “adding AI to a web app” meant exactly one thing: an HTTP call to an API, with network latency, per-token cost, and user data leaving the device. For many use cases that’s still the right call. But for a growing share of features — semantic search over local content, text classification, OCR, background removal, transcription — the model can run directly in the browser, on the user’s GPU, through WebGPU.

The result: zero network latency after the first load, zero server cost for inference, and data that never leaves the device. In exchange, you’re the one managing model downloads, memory, and compatibility across wildly different devices. Let’s see how, with real TypeScript code.

Why WebGPU changes things

WebGL already let you use the GPU from the browser, but it was an API designed for drawing triangles: doing general-purpose compute meant disguising your data as textures. WebGPU instead exposes compute shaders as first-class citizens, with memory buffers, compute pipelines, and a model much closer to Vulkan/Metal/Direct3D 12. That’s exactly the primitive the matrix multiplications at the core of every neural network need.

In practice you’ll almost never write shaders by hand: runtimes like ONNX Runtime Web and libraries like Transformers.js (which uses it under the hood) translate the model graph into WebGPU operations for you. When WebGPU isn’t available, they fall back to WebAssembly on the CPU — slower, but it works everywhere.

ℹ️ Support in 2026

WebGPU has been available in Chrome and Edge since 2023, in Safari since Safari 26, and in Firefox (first on Windows, then on other platforms). Coverage is broad but not universal — especially on Linux and older mobile GPUs — so feature detection with a fallback is mandatory, not optional.

The stack: what to pick

Library Level When to use it
@huggingface/transformers High (ready-made pipelines) Embeddings, classification, translation, OCR, speech-to-text with Hugging Face models
onnxruntime-web Mid (ONNX sessions) You have your own model exported to ONNX and want control over input/output tensors
@mlc-ai/web-llm High (OpenAI-style API) Conversational LLMs (Llama, Qwen, Phi…) entirely in the browser

For this article’s example we’ll use Transformers.js: it’s the fastest way in, and it covers the use case most apps actually benefit from — semantic search.

Goal: the user types “how do I cancel a fetch request” and finds the document titled “AbortController and cancellation signals”, even with no words in common. We need an embedding model that turns each text into a numeric vector; texts with similar meaning produce vectors that are close together.

We’ll use Xenova/all-MiniLM-L6-v2: small (about 23 MB quantized to 8 bits), fast, 384-dimensional vectors. For non-English content, a multilingual model like Xenova/paraphrase-multilingual-MiniLM-L12-v2 is a better fit, at a somewhat larger size.

npm install @huggingface/transformers
npm install -D @webgpu/types

@webgpu/types adds the type definitions for navigator.gpu, which TypeScript doesn’t yet ship in its default DOM lib:

// tsconfig.json
{
  "compilerOptions": {
    "types": ["@webgpu/types"]
  }
}

Rule number one: inference goes in a Web Worker

Loading a model means downloading tens of MB, decompressing them, compiling shaders. Even with the GPU doing the heavy lifting, tensor preparation and tokenization run in JavaScript. Doing it on the main thread means freezing the UI for hundreds of milliseconds — exactly the problem I described in the article on Web Workers, and one that shows up as a bad INP in your Core Web Vitals.

Good news: WebGPU is also available inside Dedicated Workers (navigator.gpu exists on WorkerNavigator), so everything — download, compilation, inference — can live off the main thread.

Let’s start with the message protocol, typed as a discriminated union so the main thread and the worker can’t drift apart:

// embed.protocol.ts
export type WorkerRequest =
  | { type: "load" }
  | { type: "embed"; id: number; texts: string[] };

export type WorkerResponse =
  | { type: "progress"; file: string; progress: number }
  | { type: "ready"; device: "webgpu" | "wasm" }
  | { type: "result"; id: number; vectors: Float32Array; dims: number }
  | { type: "error"; id?: number; message: string };

The worker

// embed.worker.ts
import { pipeline, type FeatureExtractionPipeline } from "@huggingface/transformers";
import type { WorkerRequest, WorkerResponse } from "./embed.protocol";

const MODEL_ID = "Xenova/all-MiniLM-L6-v2";

let extractor: Promise<FeatureExtractionPipeline> | null = null;
let device: "webgpu" | "wasm" = "wasm";

const post = (msg: WorkerResponse, transfer: Transferable[] = []) =>
  self.postMessage(msg, { transfer });

async function detectDevice(): Promise<"webgpu" | "wasm"> {
  if (!("gpu" in navigator)) return "wasm";
  const adapter = await navigator.gpu.requestAdapter();
  return adapter ? "webgpu" : "wasm";
}

function getExtractor() {
  // Singleton: one model instance, even if multiple "load" messages arrive
  extractor ??= (async () => {
    device = await detectDevice();
    const pipe = await pipeline("feature-extraction", MODEL_ID, {
      device,
      // q8 is much faster on WASM; on the GPU fp32 avoids precision issues
      dtype: device === "webgpu" ? "fp32" : "q8",
      progress_callback: (p) => {
        if (p.status === "progress") {
          post({ type: "progress", file: p.file, progress: p.progress });
        }
      },
    });
    return pipe;
  })();
  return extractor;
}

self.onmessage = async (event: MessageEvent<WorkerRequest>) => {
  const msg = event.data;
  try {
    if (msg.type === "load") {
      await getExtractor();
      post({ type: "ready", device });
      return;
    }

    if (msg.type === "embed") {
      const pipe = await getExtractor();
      const output = await pipe(msg.texts, { pooling: "mean", normalize: true });
      const vectors = output.data as Float32Array;
      const dims = output.dims[1];
      // Transfer the buffer instead of copying it: zero-copy between threads
      post({ type: "result", id: msg.id, vectors, dims }, [vectors.buffer]);
    }
  } catch (err) {
    post({
      type: "error",
      id: msg.type === "embed" ? msg.id : undefined,
      message: err instanceof Error ? err.message : String(err),
    });
  }
};

A few non-obvious choices:

  • pooling: "mean", normalize: true produces a single vector per text, already normalized to length 1. That makes cosine similarity equal to a plain dot product — cheaper to compute.
  • The buffer is transferred, not copied. With hundreds of documents at 384 floats each, skipping the structured clone matters. After postMessage, vectors in the worker is unusable (detached): that’s intended.
  • The dtype depends on the device. 8-bit quantization cuts size and time on the CPU; on the GPU the speed difference is smaller and fp32 is the most robust choice. If the adapter supports the shader-f16 feature, fp16 is a good middle ground.

The main-thread client

// embedder.ts
import type { WorkerRequest, WorkerResponse } from "./embed.protocol";

export class Embedder {
  private worker = new Worker(new URL("./embed.worker.ts", import.meta.url), {
    type: "module",
  });
  private nextId = 0;
  private pending = new Map<
    number,
    { resolve: (v: Float32Array[]) => void; reject: (e: Error) => void }
  >();
  readonly ready: Promise<"webgpu" | "wasm">;

  constructor(onProgress?: (file: string, progress: number) => void) {
    this.ready = new Promise((resolve, reject) => {
      this.worker.onmessage = (event: MessageEvent<WorkerResponse>) => {
        const msg = event.data;
        switch (msg.type) {
          case "progress":
            onProgress?.(msg.file, msg.progress);
            break;
          case "ready":
            resolve(msg.device);
            break;
          case "result": {
            const out: Float32Array[] = [];
            for (let i = 0; i < msg.vectors.length; i += msg.dims) {
              out.push(msg.vectors.subarray(i, i + msg.dims));
            }
            this.pending.get(msg.id)?.resolve(out);
            this.pending.delete(msg.id);
            break;
          }
          case "error": {
            const error = new Error(msg.message);
            if (msg.id === undefined) {
              reject(error);
            } else {
              this.pending.get(msg.id)?.reject(error);
              this.pending.delete(msg.id);
            }
          }
        }
      };
    });
    this.send({ type: "load" });
  }

  private send(msg: WorkerRequest) {
    this.worker.postMessage(msg);
  }

  async embed(texts: string[]): Promise<Float32Array[]> {
    await this.ready;
    const id = this.nextId++;
    return new Promise((resolve, reject) => {
      this.pending.set(id, { resolve, reject });
      this.send({ type: "embed", id, texts });
    });
  }

  dispose() {
    this.worker.terminate();
  }
}

new URL("./embed.worker.ts", import.meta.url) is the pattern Vite, webpack 5, and other modern bundlers recognize to emit a separate chunk for the worker — no extra configuration.

// search.ts
import { Embedder } from "./embedder";

const dot = (a: Float32Array, b: Float32Array) => {
  let sum = 0;
  for (let i = 0; i < a.length; i++) sum += a[i] * b[i];
  return sum;
};

const docs = [
  "AbortController and cancellation signals",
  "Debounce and throttle in JavaScript",
  "Content Security Policy: nonces and strict-dynamic",
  "Web Workers: moving work off the main thread",
];

const embedder = new Embedder((file, p) => console.log(`${file}: ${p.toFixed(0)}%`));
console.log("Device:", await embedder.ready);

// Computed once, then optionally persisted to IndexedDB
const docVectors = await embedder.embed(docs);

export async function search(query: string, k = 3) {
  const [q] = await embedder.embed([query]);
  return docs
    .map((text, i) => ({ text, score: dot(q, docVectors[i]) }))
    .sort((a, b) => b.score - a.score)
    .slice(0, k);
}

console.table(await search("how do I cancel a fetch request"));

For a few thousand documents, a linear scan is more than enough (a few million multiplications, under a millisecond). Beyond that, reach for an approximate index (HNSW) or move the scoring into the worker too.

💡 Tip

Document embeddings don’t change until the text does: store them in IndexedDB along with a content hash and recompute only the ones that changed. The model itself is already cached by Transformers.js through the browser’s Cache API, so from the second load the download disappears.

Handling memory and very different devices

This is where a demo and a production feature part ways. The same code will run on a workstation with 24 GB of VRAM and on a mid-range phone with shared memory. A few checks to run before downloading the model:

// capabilities.ts
export async function inspectDevice() {
  const gpu = "gpu" in navigator ? await navigator.gpu.requestAdapter() : null;

  const storage = await navigator.storage?.estimate?.();

  return {
    webgpu: !!gpu,
    // Single-buffer limit: a layer's weights must fit inside it
    maxBufferSize: gpu?.limits.maxBufferSize ?? 0,
    maxStorageBinding: gpu?.limits.maxStorageBufferBindingSize ?? 0,
    fp16: gpu?.features.has("shader-f16") ?? false,
    // Chromium only, rounded and capped at 8 GB for privacy: a hint, not a number
    deviceMemoryGB: (navigator as { deviceMemory?: number }).deviceMemory,
    // Space available for caching the model
    quotaFreeMB: storage ? ((storage.quota ?? 0) - (storage.usage ?? 0)) / 1e6 : undefined,
  };
}

With this information you can apply a tiered strategy:

  1. WebGPU + plenty of memory → larger model or fp32/fp16 precision.
  2. WebGPU with low limits, or deviceMemory ≤ 4 → small quantized model (q8/q4).
  3. No WebGPU → WASM with a quantized model, and maybe a UI hint that the operation will be slower.
  4. Not enough storage quota, or a metered connection (navigator.connection?.saveData) → don’t download anything automatically: ask the user, or fall back to a server API.

Another event to handle is GPU device loss: the driver can reset, the OS can reclaim memory, a background tab can get deprioritized. If you work with WebGPU directly, device.lost is a Promise that resolves when it happens; with Transformers.js or ONNX Runtime you’ll see it as an error during inference. Either way the strategy is the same: terminate the worker, spawn a new one, and reload the model (from cache, so no new download).

⚠ Warning

ONNX Runtime’s multi-threaded WASM fallback uses SharedArrayBuffer, which requires a cross-origin isolated page (Cross-Origin-Opener-Policy: same-origin and Cross-Origin-Embedder-Policy: require-corp headers). Without them the runtime still works, but single-threaded — much slower. Be careful: COEP can break third-party embeds and scripts that don’t send the right CORP/CORS headers.

What about LLMs?

The same approach extends to generative models with WebLLM, which exposes an OpenAI-compatible API:

import { CreateMLCEngine } from "@mlc-ai/web-llm";

const engine = await CreateMLCEngine("Qwen2.5-1.5B-Instruct-q4f16_1-MLC", {
  initProgressCallback: (p) => console.log(p.text),
});

const reply = await engine.chat.completions.create({
  messages: [{ role: "user", content: "Summarize what WebGPU is in one sentence." }],
});
console.log(reply.choices[0].message.content);

WebLLM also offers CreateWebWorkerMLCEngine to run everything in a worker behind the same API. But the proportions change: a 1-3 billion parameter model quantized to 4 bits is one or more GB to download and needs about as much GPU memory. It makes sense for tools people use repeatedly (an editor, an internal app), much less for a landing page visited once.

When it makes sense (and when it doesn’t)

Client-side inference pays off when:

  • the model is small (tens to a few hundred MB) and reused several times per session;
  • the data is sensitive and shouldn’t leave the device (documents, photos, private text);
  • you want offline or local-first features;
  • request volume would make a pay-per-use API expensive.

The server is still the better choice when you need a large, top-quality model, when target devices are old or underpowered, or when the first use is also the only one and a download of tens of MB isn’t justified.

Frequently asked questions

❓ Does WebGPU work inside a Web Worker?

Yes. navigator.gpu is also exposed on WorkerNavigator in Dedicated Workers, so download, shader compilation, and inference can all stay off the main thread.

❓ What happens if the browser doesn't support WebGPU?

Transformers.js and ONNX Runtime Web can use WebAssembly on the CPU. Pick device: "wasm" (or let the runtime fall back) and prefer a quantized model: slower, but it works everywhere.

❓ Is the model downloaded again on every visit?

No. Transformers.js stores model files in the browser’s Cache API, so from the second load they’re read locally. The browser can still evict that cache when storage runs low, so your code must always handle a fresh download.

❓ How big is a browser-friendly model?

For embeddings and classification, 20-100 MB is enough. OCR and speech-to-text models range from a few tens to a few hundred MB. LLMs, even small ones quantized to 4 bits, start at around a GB.

Conclusion

WebGPU has turned in-browser inference from an experiment into a real architectural option. The model code is the easy part — a Transformers.js pipeline is three lines. The real work is everything around it: keeping the main thread free with a Web Worker, choosing model and precision based on the device, handling cache, fallbacks, and device loss. Get that right and you ship AI features with zero latency, zero server cost, and privacy by design.

If performance is what you care about, the starting point is the article on Web Workers in JavaScript; to cleanly cancel searches still in flight while the user keeps typing, the same pattern described in AbortController applies.

Condividi:XLinkedInFacebookWhatsApp
Diego Betto

Written by

Diego Betto

Co-Founder & CTO at PAPION. Senior full-stack engineer specializing in React, TypeScript, Node.js, and application security.