Skip to content

[WebGPU] Support 2-bit quantization in GatherBlockQuantized #28895

Description

@hemanth

Describe the feature request

The GatherBlockQuantized op in the WebGPU backend currently only supports 4-bit and 8-bit quantization. Attempting to run a model with 2-bit weights fails with:

ERROR_CODE: 1, ERROR_MESSAGE: .../gather_block_quantized.h:55
onnxruntime::contrib::webgpu::GatherBlockQuantized::GatherBlockQuantized(const OpKernelInfo &)
bits_ == 4 || bits_ == 8 was false. 'bits' must be 4 or 8.

Motivation

Google recently released Gemma 4 QAT checkpoints — Quantization-Aware Training models that use 2-bit weights while maintaining near-original quality. The onnx-community has already exported these as ONNX:

These models are specifically designed for on-device / edge inference and would be a great fit for browser-based ML via WebGPU + transformers.js. However, they can't run because the WebGPU backend rejects the 2-bit weight format.

For context, I'm building Eloquent — an on-device AI dictation app that runs Gemma 4 entirely in the browser using transformers.js + WebGPU. The current q4f16 models work great, but they're ~4.9 GB for the E4B variant. The QAT 2-bit models would cut that roughly in half while preserving quality, which is a huge deal for browser-based inference where every MB matters.

What I've tried

  • Loading the QAT mobile ONNX models with dtype: "q2f16" via transformers.js → fails with the error above
  • Per-component dtype mapping (vision encoder at fp16, everything else at q2f16) → same error on the decoder

Suggested behavior

Extend GatherBlockQuantized (and any other affected ops) in the WebGPU backend to support bits == 2, similar to how 4-bit and 8-bit are handled today.

Environment

  • ONNX Runtime version: Latest via transformers.js CDN (ort-wasm-simd-threaded.jsep.wasm)
  • Browser: Chrome 131 / Edge 131 (WebGPU enabled)
  • Platform: macOS, also tested on Android Chrome
  • Model: onnx-community/gemma-4-E4B-it-qat-mobile-ONNX

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    ep:WebGPUort-web webgpu providerplatform:mobileissues related to ONNX Runtime mobile; typically submitted using templateplatform:webissues related to ONNX Runtime web; typically submitted using templatequantizationissues related to quantizationstaleissues that have not been addressed in a while; categorized by a bot

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions