Skip to content

docs(stream): clarify VAD mode timestamps and CWD model path errors - #4019

Merged
danbev merged 2 commits into
ggml-org:masterfrom
apollo-2006:patch-1
Sep 18, 2026
Merged

danbev merged 2 commits into
ggml-org:masterfrom
apollo-2006:patch-1

Conversation

@apollo-2006

Copy link
Copy Markdown
Contributor

Added documentation for undocumented timestamp behavior in VAD mode and clarified CWD errors to prevent parser breaks," and hit submit

AI use: I found this while building against the library and used an AI assistant to help verify the relevant source (file/line refs above). The report and the documentation wording are mine, and I've checked every claim against the source myself.

Copilot AI lite review requested due to automatic review settings August 27, 2026 00:54

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR updates the whisper-stream example documentation to make its stdout output (especially in VAD / --step 0 mode) easier to parse reliably, and to clarify how model path resolution depends on the process working directory.

Changes:

  • Documented the two distinct stdout output shapes for --step > 0 vs --step 0 (VAD mode), including the VAD block markers and timestamped segment lines.
  • Clarified that -m/--model is resolved relative to the process CWD and that CWD mistakes can surface as the generic “failed to initialize whisper context” error.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread examples/stream/README.md Outdated
Comment thread examples/stream/README.md
@apollo-2006

Copy link
Copy Markdown
Contributor Author

@danbev - you wrote #3065, so you're the right person to ask about the part I couldn't settle while writing this up.

When VAD is on, the timestamps that come back are relative to the filtered audio and get mapped back onto the original timeline. Reading the review thread on #3065, the mapping scales speech back across the original segment length rather than preserving where the silence actually sat. For a caller doing word-level alignment, is that difference meant to be treated as noise, or is preserving true silence positions something you'd want a follow-up for?

I hit this building a local voice assistant with a "Delphi" wake word on top of whisper-stream, so my failure mode was practical rather than theoretical: I was discarding timestamped lines as log noise until I realised the prefix was the data.

I'm a third-year CS student and I'm using this project mostly as an excuse to get better at reading code I didn't write. Thanks for #3065 either way.

@danbev

danbev commented Aug 28, 2026

Copy link
Copy Markdown
Member

@apollo-2006 The stream example uses vad_simple and not the VAD implementation from #3065.

For your question, please take a look at #3910 which may be relevant.

@apollo-2006

Copy link
Copy Markdown
Contributor Author

Thanks, that's the part I had wrong. I was looking in the wrong place.

Checked stream.cpp: use_vad is just n_samples_step <= 0, and the only VAD is vad_simple in examples/common.cpp, which is an energy gate returning a bool. It never cuts silence or emits segments, and wparams.vad is never set, so nothing from #3065 runs here. My question didn't apply.

#3910 does answer what I was actually asking. Snapping a token that lands in removed silence to the nearer boundary, instead of interpolating across the gap, is the behaviour I meant. Merged, so nothing needed there.

One thing I did find: in --step 0, the ### Transcription N t0/t1 are wall-clock since stream start, while the [hh:mm:ss -->] prefixes are relative to the chunk passed to whisper_full. Two clocks, and a parser that mixes them gets it wrong. Want me to add that to this PR? Seems more useful than what's in there now.

@apollo-2006
apollo-2006 marked this pull request as draft September 7, 2026 16:37
AI use: I found this while building against the library and used an AI
assistant to help verify the relevant source (file/line refs
above). The report and the documentation wording are mine, and I've checked every claim against the source myself.
@apollo-2006
apollo-2006 marked this pull request as ready for review September 8, 2026 16:23
@apollo-2006

Copy link
Copy Markdown
Contributor Author

@danbev this is out of draft and ready for review. Force-pushing does not notify, so flagging it here.

One thing I noticed rereading the sample output: the block prints ### Transcription 0 START | t0 = 0 ms | t1 = 4000 ms two lines above [00:00:00.000 --> 00:00:03.480]. Those are two different clocks sitting next to each other, and nothing on the page says which is which.

Worth a sentence distinguishing them here, or is that out of scope for a stream README?

Comment thread examples/stream/README.md Outdated
Co-authored-by: Daniel Bevenius <daniel.bevenius@gmail.com>
@apollo-2006
apollo-2006 requested a lite review from Copilot September 13, 2026 03:15

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@danbev
danbev merged commit d5d6e59 into ggml-org:master Sep 18, 2026
40 checks passed
bygreencn added a commit to bygreencn/whisper.cpp that referenced this pull request Sep 23, 2026
* ggerganov/master: (81 commits)
  fix(yt-wsp): Resolve script path without GNU realpath (ggml-org#4072)
  cli : load backends after validating input files (ggml-org#4069)
  ci : update android-actions to v4.0.4 (ggml-org#4074)
  docs : clarify VAD mode timestamps and CWD model path errors (ggml-org#4019)
  readme : document the ANEForge encoder backend (ggml-org#4073)
  whisper : optional ANEForge encoder backend (Apple Neural Engine) (ggml-org#3905)
  whisper : fix int overflow in whisper_full_parallel chunk offsets (ggml-org#4044)
  sync : ggml
  ggml : bump version to 0.24.0 (ggml/1627)
  tests(s390x): add non-vxe build to tests (llama/28776)
  sycl: rfc: Use radix select for top_k (llama/28670)
  ggml-cpu : disable PCH and fix CACHE_LINE_SIZE ambiguity to fix heap corruption (llama/28882)
  sycl : fix oneDNN scratchpad breaking the pool free order (llama/28704)
  ggml-cuda: fallback to F32 on device without BF16 hardware acceleration (llama/28846)
  ggml-cpu(s390x): guard VXE-only repack helpers (llama/28775)
  sycl : Fix get mem error (llama/28227)
  vulkan: workaround NV queuesubmit driver bug (llama/28830)
  opencl: apply the noshuffle row-alignment rule to q4_K, q5_K and q8_0, not just q6_K (llama/28575)
  ggml-cuda: hip add specific config table for AMD GCN (llama/27841)
  syscl : Handle (fail gracefully) unsupported tq1_0 quants (llama/28681)
  ...
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants