Reported by mail by Max Rempel (Pixel 7 Pro, Android 17, 6.2.0 from Play), who read the source and
arrived at the mechanism himself. The analysis below is his, verified against the tree; the scope
and the implementation notes are mine.
Symptom. In realtime mode the field cannot be corrected while dictation runs. Delete a word the
model got wrong — noise in the room, someone else talking — and the next update puts it back. Worse
at length: select all, delete, and the next update pastes the entire conversation so far back into
the field. The user ends up with a field full of text they had already removed, and no way to fix
anything without stopping first.
Mechanism. DictateController keeps realtimeShown, its own copy of what it believes is in the
field, and every write is computed against that copy rather than against the field:
setDictationPreview(full, realtimeShown), commitDictationFinal(outputText, realtimeShown),
clearDictationPreview(realtimeShown).
All three land in ImeDictationSink and none of them reads the field first:
applyDictationDiff (DictationSink.kt:143) keeps the common prefix with the remembered string
and calls replaceTextBeforeCursor(old.length - cp, …) for the divergent tail.
commitDictationFinal (DictationSink.kt:126) does the same against prevText.
clearDictationPreview (DictationSink.kt:134) deletes prevText.length characters before the
cursor outright.
So the moment the user edits by hand, the remembered string stops describing reality. The delete
count is then measured against text that is no longer there — it eats the user's own words — and the
tail that "diverges" is the whole transcript, which gets committed again.
The fix has a precedent in the same file. deleteLastText (DictationSink.kt:112) already
refuses to act when the field does not match: it only undoes when textBeforeSelection ends with
exactly what was inserted. The three realtime paths should make the same comparison and, when it
fails, accept the field as the new baseline instead of trying to reconcile against a copy that
is already wrong. Deleted text then stays deleted and the next words are appended at the cursor,
which is what the user expects.
The trap in doing that. EditorContent.textBeforeSelection is capped at 256 characters
(NumCharsBeforeCursor, AbstractEditorInstance.kt:56). realtimeShown routinely exceeds that in a
long dictation, so a plain endsWith(realtimeShown) would fail on length alone and re-baseline on
every single update — turning a correctness fix into a permanent one. The comparison has to be
bounded to the available window: compare the last n characters, where n is the smaller of the
window and the remembered string, and treat "the window is shorter than what we remember" as a match
rather than a mismatch.
Scope. ImeDictationSink only. AccessibilitySink ignores prevText and manages its own
preview through the service, and RecognitionSink has no field access at all — neither can drift
this way.
Worth deciding while in here: whether a re-baseline should also reset the realtime provider's
notion of the session, or only ours. Appending at the cursor with a provider that still holds the
full transcript means the next revision it sends can reach back across the user's edit again. My
inclination is that a re-baseline truncates what we will ever revise — anything before the edit
point is the user's now, not ours — but that needs a device test against a provider that revises
aggressively, not an argument.
Reported by mail by Max Rempel (Pixel 7 Pro, Android 17, 6.2.0 from Play), who read the source and
arrived at the mechanism himself. The analysis below is his, verified against the tree; the scope
and the implementation notes are mine.
Symptom. In realtime mode the field cannot be corrected while dictation runs. Delete a word the
model got wrong — noise in the room, someone else talking — and the next update puts it back. Worse
at length: select all, delete, and the next update pastes the entire conversation so far back into
the field. The user ends up with a field full of text they had already removed, and no way to fix
anything without stopping first.
Mechanism.
DictateControllerkeepsrealtimeShown, its own copy of what it believes is in thefield, and every write is computed against that copy rather than against the field:
setDictationPreview(full, realtimeShown),commitDictationFinal(outputText, realtimeShown),clearDictationPreview(realtimeShown).All three land in
ImeDictationSinkand none of them reads the field first:applyDictationDiff(DictationSink.kt:143) keeps the common prefix with the remembered stringand calls
replaceTextBeforeCursor(old.length - cp, …)for the divergent tail.commitDictationFinal(DictationSink.kt:126) does the same againstprevText.clearDictationPreview(DictationSink.kt:134) deletesprevText.lengthcharacters before thecursor outright.
So the moment the user edits by hand, the remembered string stops describing reality. The delete
count is then measured against text that is no longer there — it eats the user's own words — and the
tail that "diverges" is the whole transcript, which gets committed again.
The fix has a precedent in the same file.
deleteLastText(DictationSink.kt:112) alreadyrefuses to act when the field does not match: it only undoes when
textBeforeSelectionends withexactly what was inserted. The three realtime paths should make the same comparison and, when it
fails, accept the field as the new baseline instead of trying to reconcile against a copy that
is already wrong. Deleted text then stays deleted and the next words are appended at the cursor,
which is what the user expects.
The trap in doing that.
EditorContent.textBeforeSelectionis capped at 256 characters(
NumCharsBeforeCursor,AbstractEditorInstance.kt:56).realtimeShownroutinely exceeds that in along dictation, so a plain
endsWith(realtimeShown)would fail on length alone and re-baseline onevery single update — turning a correctness fix into a permanent one. The comparison has to be
bounded to the available window: compare the last n characters, where n is the smaller of the
window and the remembered string, and treat "the window is shorter than what we remember" as a match
rather than a mismatch.
Scope.
ImeDictationSinkonly.AccessibilitySinkignoresprevTextand manages its ownpreview through the service, and
RecognitionSinkhas no field access at all — neither can driftthis way.
Worth deciding while in here: whether a re-baseline should also reset the realtime provider's
notion of the session, or only ours. Appending at the cursor with a provider that still holds the
full transcript means the next revision it sends can reach back across the user's edit again. My
inclination is that a re-baseline truncates what we will ever revise — anything before the edit
point is the user's now, not ours — but that needs a device test against a provider that revises
aggressively, not an argument.