Skip to content

whisper : fix int overflow in whisper_full_parallel chunk offsets - #4044

Merged
danbev merged 1 commit into
ggml-org:masterfrom
kmadiar:fix-parallel-offset-overflow
Sep 15, 2026
Merged

danbev merged 1 commit into
ggml-org:masterfrom
kmadiar:fix-parallel-offset-overflow

Conversation

@kmadiar

@kmadiar kmadiar commented Sep 7, 2026

Copy link
Copy Markdown
Contributor

Fixes #4039.

whisper_full_parallel() computed the per-chunk time offset as 100 * ((i + 1) * n_samples_per_processor) / WHISPER_SAMPLE_RATE in int before adding it to the int64_t offset_t. For chunk boundaries later than ~22 min at 16 kHz (100 * samples > INT_MAX) the multiplication overflows, corrupting the merged segment t0/t1 and the split times printed in the log.

This PR computes the offset once per chunk in int64_t (100LL * ...) and reuses it for both the segment merge and the split-time log.

Verification — 46 min file (samples/jfk.wav looped), --processors 2, boundary at 23:00:

Before, built with -fsanitize=undefined:

src/whisper.cpp:7980:30: runtime error: signed integer overflow: 100 * 22080000 cannot be represented in type 'int'
src/whisper.cpp:7981:30: runtime error: signed integer overflow: 100 * 22080000 cannot be represented in type 'int'
src/whisper.cpp:8023:9:  runtime error: signed integer overflow: 100 * 22080000 cannot be represented in type 'int'
whisper_full_parallel: split 1 - 00:-21:-44.-350

After: no overflow reports, split 1 - 00:23:00.000, all segment offsets in the JSON output non-negative and monotonic. --processors 4 (Release) reports splits at 00:11:30 / 00:23:00 / 00:34:30.

Note: on Apple clang Release builds the UB happens to yield the correct value, which is why this is mostly seen on MSVC/Windows (as in the issue); the UBSan build reproduces it on macOS.

I have a small regression test (whisper_full_parallel on a 45 min silent buffer with duration_ms = 1000, checking the logged split time via whisper_log_set; ~3 s, fails on master with 00:-22:-14.-350). Happy to add it to this PR or a follow-up if you'd like it.

Token-level timestamps in parallel mode (#2036, #3726) are a separate issue and unchanged here.

@danbev
danbev merged commit da54572 into ggml-org:master Sep 15, 2026
43 of 47 checks passed
bygreencn added a commit to bygreencn/whisper.cpp that referenced this pull request Sep 23, 2026
* ggerganov/master: (81 commits)
  fix(yt-wsp): Resolve script path without GNU realpath (ggml-org#4072)
  cli : load backends after validating input files (ggml-org#4069)
  ci : update android-actions to v4.0.4 (ggml-org#4074)
  docs : clarify VAD mode timestamps and CWD model path errors (ggml-org#4019)
  readme : document the ANEForge encoder backend (ggml-org#4073)
  whisper : optional ANEForge encoder backend (Apple Neural Engine) (ggml-org#3905)
  whisper : fix int overflow in whisper_full_parallel chunk offsets (ggml-org#4044)
  sync : ggml
  ggml : bump version to 0.24.0 (ggml/1627)
  tests(s390x): add non-vxe build to tests (llama/28776)
  sycl: rfc: Use radix select for top_k (llama/28670)
  ggml-cpu : disable PCH and fix CACHE_LINE_SIZE ambiguity to fix heap corruption (llama/28882)
  sycl : fix oneDNN scratchpad breaking the pool free order (llama/28704)
  ggml-cuda: fallback to F32 on device without BF16 hardware acceleration (llama/28846)
  ggml-cpu(s390x): guard VXE-only repack helpers (llama/28775)
  sycl : Fix get mem error (llama/28227)
  vulkan: workaround NV queuesubmit driver bug (llama/28830)
  opencl: apply the noshuffle row-alignment rule to q4_K, q5_K and q8_0, not just q6_K (llama/28575)
  ggml-cuda: hip add specific config table for AMD GCN (llama/27841)
  syscl : Handle (fail gracefully) unsupported tq1_0 quants (llama/28681)
  ...
rmorse added a commit to operator-kit/whisper-cpp-plus-rs that referenced this pull request Sep 24, 2026
* chore: update whisper.cpp pin to 1.9.4-dev (stream-pcm de8fb5fd)

Moves the pinned rmorse/whisper.cpp stream-pcm fork from ddfe1196
(v1.8.6) to de8fb5fd (tag v1.9.4-dev-stream-pcm), based on upstream
master after v1.9.3. whisper.h changes in this range are additive only,
so no binding or API changes are needed. Parakeet and the new VAD
timestamp/segment accessors are left for a follow-up feature branch.

Picks up the upstream whisper_full_parallel timestamp overflow fix
(ggml-org/whisper.cpp#4044).

* docs: correct changelog for whisper.cpp 1.9 pin

Drop the whisper_full_parallel fix entry: WhisperState::full_parallel reads results from its own state while upstream writes them to the context's default state, so the upstream fix never reached crate users (to be fixed separately). Document user-visible upstream behaviour changes in the pinned range and why the new VAD accessors are not exposed.

* docs: recommend the crate's VAD over whisper.cpp's built-in VAD

whisper.cpp's built-in whisper_full VAD does not run for whisper_full_with_state, which the crate uses for all transcription (ggml-org/whisper.cpp#3423). Point users at WhisperVadProcessor, EnhancedWhisperVadProcessor and WhisperStreamPcm::with_vad instead.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

whisper_full_parallel: signed 32-bit overflow corrupts segment timestamps for long audio

2 participants