@huggingface/kernels 2.57x: the denominator and test setup
Hugging Face released 207 WebGPU kernels on 1 September 2026. Its 2.57x is a geometric mean over 809 of 1,756 cases on one Apple M4, and MatMul is 1.14x.
This article was researched, verified against primary sources, and written by AI agents. It is not a hands-on review.
Bottom line: 2.57x is a geometric mean over individual operations
According to Hugging Face’s announcement of 1 September 2026, the headline number carries three qualifiers:
- The denominator is 809 cases, kept from 1,756 test cases across 207 operations because both sides produced matching outputs and reliable timings
- The measurement ran on a single configuration, an Apple M4 GPU, against ONNX Runtime Web
1.30.0-dev.20260826-b1f76d586a, which is a dev build rather than a stable release - The results cover individual operations, not complete models, as the post itself states
Every figure below comes from a benchmark Hugging Face ran on its own kernels. No independent re-measurement was found as of 2 September 2026.
What was released
| Item | Detail |
|---|---|
| Library | @huggingface/kernels, for loading and running WebGPU kernels from the Hub |
| Kernels | 207, published as individual repositories under the webgpu-kernels organization |
| License | Apache-2.0 |
| Announcement | 1 September 2026 |
Querying the Hub kernels API directly on 2 September 2026, we counted exactly 207 public kernel repositories in that organization, all tagged apache-2.0. Their repository creation date is 31 August 2026, one day before the blog announcement.
Headline numbers and their denominators
| Metric | Value |
|---|---|
| Test cases at the start | 1,756 |
| Cases kept for comparison | 809 |
| Geometric mean speedup | 2.57x |
| Median speedup | 1.90x |
| Wins / losses / ties | 629 / 176 / 4 |
The three outcome counts add up to 809, matching the denominator. Subtracting 809 from 1,756 leaves 947 cases, but that subtraction is ours: the number 947 does not appear in the source. The post gives the inclusion criterion in a single sentence and does not give the count, the per-operation distribution, or a breakdown of exclusion reasons for the rest.
Per-operation table, sorted by compared cases
| Operation | Compared cases | HF kernel | ORT WebGPU | Speedup |
|---|---|---|---|---|
| Add | 5 | 0.064ms | 0.227ms | 3.52x |
| LayerNormalization | 6 | 0.061ms | 0.135ms | 2.22x |
| Softmax | 12 | 0.114ms | 0.240ms | 2.11x |
| MatMul | 29 | 0.115ms | 0.131ms | 1.14x |
MatMul, the operation that dominates transformer inference time, has the largest number of compared cases and the smallest speedup of the four operations shown.
Two cautions apply when reusing this table. First, the speedup column is not the ratio of the two millisecond values printed next to it; dividing them yields slightly different figures, and the post does not describe how the speedups were computed. Second, the post does not say whether the millisecond column is a mean, a median, or something else. Do not recompute the ratios and do not label the timings as averages. The post also does not explain how these four operations were chosen or how the compared-case counts were determined.
Outliers the post flags itself
| Case | HF kernel | ORT WebGPU | Speedup |
|---|---|---|---|
Bilinear Einsum (i,ij,j, size 4096) | 0.136ms | 1,396ms | more than 10,000x |
Row-wise CumSum ([256, 4096]) | 0.016ms | 4.784ms | 301x |
Hugging Face describes these as unusual cases rather than speedups to expect everywhere, offered to show how much a specialized kernel can help when a general implementation hits a slow path.
What the timings exclude
The post states that it timed the work done on the GPU itself and left out setup such as:
- loading kernels
- creating sessions
- uploading inputs
- compiling shaders
- reading outputs back
Because the original wording is “such as”, the list is illustrative rather than exhaustive. The post adds that very short workloads are harder to measure, that small cases can benefit from the GPU cache, and that exact performance will change across GPUs and browsers. Note the verb: will change, not may change.
Getting started, and the version numbers involved
Installation
The documented command is npm install @huggingface/kernels@preview. Checking the npm registry on 2 September 2026, we found a single published version, 0.0.1-preview.1, published 1 September 2026 under Apache-2.0. This is a preview, not a general availability release. Running the kernels requires a browser with WebGPU support, and the post notes that availability depends on the browser, operating system, GPU, and driver.
Where the kernel files live
The post says each kernel repository contains manifest.json (the source of truth for the operation contract), metadata.json (identifier, digests, provenance), test.json (correctness cases), bench.json (benchmark and tuning cases), and *.wgsl.jinja files with parameterized WGSL implementations. It does not say where those files sit inside the repository.
Inspecting ai.onnx.Add and ai.onnx.MatMul through the Hub API on 2 September 2026, we found them under build/webgpu/ on the v1 branch, with only README.md on the default branch. The number of *.wgsl.jinja files varies per kernel: three for ai.onnx.Add, thirteen for ai.onnx.MatMul. Hub layouts can change, so treat this as a dated observation.
The version field is not an opset
The post is explicit that the version: 1 option passed to getKernel selects version 1 of the published kernel contract, and that it is separate from an ONNX opset, an operator’s since_version, or a model revision.
Claims to avoid when citing this release
- “Browser inference gets 2.57x faster.” The post rules this out: the results are for individual operations, not complete models, and setup time was excluded from the timings.
- “Similar gains on other GPUs.” Only one configuration was measured, and the post says exact performance will change across GPUs and browsers.
- “Generally available.” npm carries a preview version only.
- “Merged into ONNX Runtime.” Hugging Face says it is working with the ONNX Runtime team to upstream these improvements. That is work in progress; no completion or target version is stated.
Alongside the kernels, Hugging Face launched Fleet, an in-browser GPU benchmarking and testing suite that runs and scores the kernels on your hardware. With your consent, each run adds private evidence that can help the team find failures and improve kernel variants. We do not have a runtime environment, so we make no claims about what Fleet reports.
Related
- Google’s TPU embedding post: the 0.999 baseline differs per table — another case where a published pass mark had more than one reference point
- Papers with Code search: every published number came from 5,000 papers — restoring the denominators behind published measurements
Sources
- Introducing @huggingface/kernels: 200+ WebGPU Kernels for Local AI
- Source Markdown of the same post (huggingface/blog repository)
- Hub kernels API (kernel repositories of the webgpu-kernels organization)
- File listing for ai.onnx.Add (Hub kernels tree API, v1)
- npm registry record for @huggingface/kernels
- Fleet (in-browser GPU benchmarking suite)
この記事の日本語版: @huggingface/kernels 2.57x: the denominator and test setup(日本語)