fix(ads): the writer's prompt names real voices, normalises the reply, and budgets each attempt on its own (#696) - #697
Merged
Conversation
… before the validator, and budgets each attempt on its own First contact on the demo (gh-#696): two ticks, zero usable spots. The prompt's abstract "TAG: <line>" placeholder was copied verbatim by the reference station's 3B model; the re-ask then shared one Llm:TimeoutSeconds budget with a 49-second first attempt and timed out by construction. - AdScriptPromptBuilder spells the example with the real tags (the CrosstalkPromptBuilder shape) and states the tag grammar in words — 79% raw pass on llama3.2:3b vs 0-29% - ApplyLineAwareHygiene folds the model's shape quirks before the fail-closed validator sees the script (quotes, tag case/spacing/parentheticals, beat labels as speakers, split tag/text lines, title and direction lines, a cast with no ANNOUNCER) — 88-92% after normalising; owner text never passes through it - each attempt gets its own timeout budget: this writer is off the air clock - MockCompletionsServer.ServeDelayMs: a slow success for the budget spec Closes #696
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to subscribe to this conversation on GitHub.
Already have an account?
Sign in.
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #696. First-contact fix for the ad writer, three halves, each pinned; bench evidence on the issue.
🔧 What changed
AdScriptPromptBuilder):"ANNOUNCER: <line>"for the announcer,"VOICE1: <line>"or"VOICE2: <line>"for up to two other voices, plus one sentence of tag grammar in words ("the word before the colon is always the VOICE speaking … no quotes, no parentheses, no stage directions; never a beat name … never the brand name"). The abstract"TAG: <line>"placeholder is gone — the reference model copied it verbatim, quotes included. This is theCrosstalkPromptBuildershape, which is why that writer never had the problem.AdScriptWriter.ApplyLineAwareHygiene, LLM path only — owner text stays verbatim per F160.4): wrapping quotes/bullets/emphasis stripped from the whole line before the tag is split; the tag folded to[A-Za-z0-9]upper-case with parentheticals dropped (Announcer,VOICE 1,PRUETT'S,LARRY (YELLING)); a beat label used as the speaker (HOOK:,Tagline:, the literalTAG:) is the announcer; an untagged line joins the previous voice (filling a bareTAG:when the model put the words on the next line); title,[bracketed]/(parenthesised)and#lines are dropped; and when no line is tagged ANNOUNCER, the most frequent voice is. Never the chat-preamble heuristic on a tag — the T400 review's hazard stays closed. A still-empty tag stays visible for the validator, as before.Llm:TimeoutSecondsbudget (NewAttemptBudget): the on-air writers share one across a re-ask because the break is imminent; this writer runs off the air clock, and on the demo's CPU-bound model one completion took 49 s of the shared 90 s, so the re-ask timed out by construction.📊 Why this shape (llama3.2:3b, the demo's model — full tables on #696)
Residuals after both are real content for the re-ask: too long, too short, more than three voices.
🧭 Two design calls inside, flagged for the ruling
AdRenderService, null voice plan), so nothing audible changes; it turns a whole re-ask into a rename. Drop the rule if you'd rather the re-ask carry it — it is one block.✅ Verified
dotnet build GenWave.sln: 0 warnings, 0 errors.ServeDelayMsknob (a slow success, distinct from the hang mode) for the budget spec.📻 What the demo did meanwhile
Four ticks, four distinct shape quirks, zero usable spots — every one is a row in the normaliser's table:
"TAG(quotes included)"TAGLARRY (YELLING)→ failed rowVOICE OVER 1ANNOUNCERalone on a line, words on the nextPatch-release candidate: the demo stays at zero usable spots per tick until this lands.