vad : reject n_encoder_layers other than 4 in model load - #4064
Merged
danbev merged 1 commit intoSep 22, 2026
Merged
Conversation
danbev
reviewed
Sep 21, 2026
n_encoder_layers is read from the VAD model file and three channel and kernel arrays are sized to it, but the encoder graph is hardcoded to four layers and indexes those arrays at [0..3]. A value below 4 reads past the allocation, so reject anything other than 4 at the point of the read.
apollo-2006
force-pushed
the
vad-reject-invalid-encoder-layers
branch
from
September 21, 2026 15:23
a894528 to
987cad1
Compare
Contributor
Author
|
comments are gone, and ive trimmed the description. thank you for the review. sorry for the inconvenience |
danbev
approved these changes
Sep 21, 2026
bygreencn
added a commit
to bygreencn/whisper.cpp
that referenced
this pull request
Sep 23, 2026
* ggerganov/master: vad : reject n_encoder_layers other than 4 in model load (ggml-org#4064) devops : reduce Vulkan Docker image to 39.4% of its original size (now 692 MB) (ggml-org#4038)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A VAD model whose header declares
n_encoder_layersother than 4 causes an out-of-bounds heap read inwhisper_vad_init_with_params, reachable from the ordinarywhisper-cli --vad -vm <file>path on any user-supplied VAD model.n_encoder_layersis read from the file and used, unbounded, to size three heap arrays, while the encoder graph below is hardcoded to four layers and indexes those arrays at[0..3]regardless of what the file declared. A value below 4 reads past the allocation.The fix rejects it at the point of the read, matching the
n_dimsguard in #3957. The graph only implements four layers, so any other value is out of range rather than something to accommodate. For a well-formed model this is a no-op.Tested with a crafted 3-layer model: SIGABRT before, rejected with exit 10 after, and heap-buffer-overflow under ASan before, clean after. The real silero v6.2.0 model transcribes unchanged.
test-vadandtest-vad-fullpass.