Skip to content

webgpu: align tensor bindings to the type block size - #28382

Merged
ServeurpersoCom merged 1 commit into
ggml-org:masterfrom
ServeurpersoCom:webgpu-get-rows-misalign
Sep 12, 2026
Merged

ServeurpersoCom merged 1 commit into
ggml-org:masterfrom
ServeurpersoCom:webgpu-get-rows-misalign

Conversation

@ServeurpersoCom

@ServeurpersoCom ServeurpersoCom commented Sep 4, 2026 •

Copy link
Copy Markdown
Contributor

Overview

This fixes the GET_ROWS failures that #28253 exposes on the WebGPU runners, the binding offset is now walked back to a multiple of the type block size so block quantized views get a valid element offset in the shader, validated locally against the CI Dawn build on an RTX PRO 6000 (207/207 GET_ROWS, 12011/12011 full suite).

Additional information

Walk the binding offset back until the distance to the tensor is a whole number of blocks, so block quantized views get a valid element offset in the shader.

Test

Tested with the CI Dawn build on an RTX PRO 6000, test-backend-ops goes from 157/207 to 207/207 on GET_ROWS and the full suite passes 12011/12011.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES Fable 5.1 MCP loop on rootless container with Nvidia GPU

@ServeurpersoCom
ServeurpersoCom requested a review from a team as a code owner September 4, 2026 11:27
@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning WebGPU labels Sep 4, 2026

@yomaytk yomaytk left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thanks!

Walk the binding offset back until the distance to the tensor is a
whole number of blocks, so block quantized views get a valid element
offset in the shader.
@ServeurpersoCom
ServeurpersoCom force-pushed the webgpu-get-rows-misalign branch from 0df5b66 to b3fb316 Compare September 7, 2026 10:25
@ServeurpersoCom

Copy link
Copy Markdown
Contributor Author

Rebase only

@yomaytk

yomaytk commented Sep 11, 2026

Copy link
Copy Markdown
Member

@reeselevine gentle ping, can you take a look at this when you have a chance?

@reeselevine

Copy link
Copy Markdown
Contributor

looks like it may need some formatting fixes?

@ServeurpersoCom

Copy link
Copy Markdown
Contributor Author

looks like it may need some formatting fixes?

Yes. The format job is failing on the base, not on this diff: the branch sits on 0cae430 where the GET_ROWS case block is still unformatted, and master already fixes it in c0b1871, so merging clears the check.

@ServeurpersoCom
ServeurpersoCom merged commit 3f5e94d into ggml-org:master Sep 12, 2026
25 of 29 checks passed
@BrewTestBot BrewTestBot mentioned this pull request Sep 14, 2026
1 task done
pl752 pushed a commit to pl752/llama.cpp that referenced this pull request Sep 15, 2026
Walk the binding offset back until the distance to the tensor is a
whole number of blocks, so block quantized views get a valid element
offset in the shader.
quimmedes pushed a commit to quimmedes/cafe-llama.cpp that referenced this pull request Sep 16, 2026
Walk the binding offset back until the distance to the tensor is a
whole number of blocks, so block quantized views get a valid element
offset in the shader.
zsogitbe pushed a commit to zsogitbe/llama.cpp that referenced this pull request Sep 17, 2026
Walk the binding offset back until the distance to the tensor is a
whole number of blocks, so block quantized views get a valid element
offset in the shader.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning WebGPU

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants