Skip to content

fix(ads): the writer's prompt names real voices, normalises the reply, and budgets each attempt on its own (#696) - #697

Merged
genwave-radio merged 1 commit into
mainfrom
fix/ads-first-contact-prompt-and-budget
Sep 5, 2026
Merged

genwave-radio merged 1 commit into
mainfrom
fix/ads-first-contact-prompt-and-budget

Conversation

@genwave-radio

@genwave-radio genwave-radio commented Sep 5, 2026 •

Copy link
Copy Markdown
Collaborator

Closes #696. First-contact fix for the ad writer, three halves, each pinned; bench evidence on the issue.

🔧 What changed

  1. The prompt spells its example with the real tags (AdScriptPromptBuilder): "ANNOUNCER: <line>" for the announcer, "VOICE1: <line>" or "VOICE2: <line>" for up to two other voices, plus one sentence of tag grammar in words ("the word before the colon is always the VOICE speaking … no quotes, no parentheses, no stage directions; never a beat name … never the brand name"). The abstract "TAG: <line>" placeholder is gone — the reference model copied it verbatim, quotes included. This is the CrosstalkPromptBuilder shape, which is why that writer never had the problem.
  2. A normalising pass before the validator (AdScriptWriter.ApplyLineAwareHygiene, LLM path only — owner text stays verbatim per F160.4): wrapping quotes/bullets/emphasis stripped from the whole line before the tag is split; the tag folded to [A-Za-z0-9] upper-case with parentheticals dropped (Announcer, VOICE 1, PRUETT'S, LARRY (YELLING)); a beat label used as the speaker (HOOK:, Tagline:, the literal TAG:) is the announcer; an untagged line joins the previous voice (filling a bare TAG: when the model put the words on the next line); title, [bracketed]/(parenthesised) and # lines are dropped; and when no line is tagged ANNOUNCER, the most frequent voice is. Never the chat-preamble heuristic on a tag — the T400 review's hazard stays closed. A still-empty tag stays visible for the validator, as before.
  3. Each attempt gets its own Llm:TimeoutSeconds budget (NewAttemptBudget): the on-air writers share one across a re-ask because the break is imminent; this writer runs off the air clock, and on the demo's CPU-bound model one completion took 49 s of the shared 90 s, so the re-ask timed out by construction.

📊 Why this shape (llama3.2:3b, the demo's model — full tables on #696)

raw pass after normalising
current prompt 0–29 % across two samples 92 %
this prompt 79 % 88 %

Residuals after both are real content for the re-ask: too long, too short, more than three voices.

🧭 Two design calls inside, flagged for the ruling

  • The lead voice becomes ANNOUNCER when the model cast nobody as one. For a generated spot every tag already maps to the station voice (AdRenderService, null voice plan), so nothing audible changes; it turns a whole re-ask into a rename. Drop the rule if you'd rather the re-ask carry it — it is one block.
  • A leading untagged line is dropped (a title, a chat preamble) rather than left for the format rule. Same trade: 3B models title their work; the re-ask was paying a full completion to delete a title.

✅ Verified

  • dotnet build GenWave.sln: 0 warnings, 0 errors.
  • Tts 857 passed (11 new: prompt shape, 9 normalisation rules, per-attempt budget — the budget spec proven discriminating by mutation: shared budget → exactly one red), Ads 98, Architecture 135. Host 2565 passed with one red — the real-Postgres migration probe, which ran while I was removing the bench's ollama container; it passed on an isolated rerun (1 m 23 s).
  • Mock server gains a ServeDelayMs knob (a slow success, distinct from the hang mode) for the budget spec.

📻 What the demo did meanwhile

Four ticks, four distinct shape quirks, zero usable spots — every one is a row in the normaliser's table:

tick first draft re-ask
1 tag "TAG (quotes included) timed out (shared budget)
2 tag "TAG LARRY (YELLING) → failed row
3 VOICE OVER 1 timed out (shared budget)
4 ANNOUNCER alone on a line, words on the next an untagged line → failed row

Patch-release candidate: the demo stays at zero usable spots per tick until this lands.

… before the validator, and budgets each attempt on its own

First contact on the demo (gh-#696): two ticks, zero usable spots. The prompt's abstract
"TAG: <line>" placeholder was copied verbatim by the reference station's 3B model; the
re-ask then shared one Llm:TimeoutSeconds budget with a 49-second first attempt and timed
out by construction.

- AdScriptPromptBuilder spells the example with the real tags (the CrosstalkPromptBuilder
  shape) and states the tag grammar in words — 79% raw pass on llama3.2:3b vs 0-29%
- ApplyLineAwareHygiene folds the model's shape quirks before the fail-closed validator
  sees the script (quotes, tag case/spacing/parentheticals, beat labels as speakers,
  split tag/text lines, title and direction lines, a cast with no ANNOUNCER) — 88-92%
  after normalising; owner text never passes through it
- each attempt gets its own timeout budget: this writer is off the air clock
- MockCompletionsServer.ServeDelayMs: a slow success for the budget spec

Closes #696
@genwave-radio genwave-radio added the bug Something isn't working label Sep 5, 2026
@genwave-radio
genwave-radio merged commit 12fd02e into main Sep 5, 2026
11 checks passed
@genwave-radio
genwave-radio deleted the fix/ads-first-contact-prompt-and-budget branch September 5, 2026 18:27
@github-actions github-actions Bot locked and limited conversation to collaborators Sep 5, 2026
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Ads first contact: the prompt's literal "TAG" is copied by the 3B model, and the re-ask shares the timeout budget — zero usable spots on the demo

1 participant