Skip to content

Realtime dictation writes against a remembered copy of the field, so anything the user deletes mid-dictation comes back #421

Description

@DevEmperor

Reported by mail by Max Rempel (Pixel 7 Pro, Android 17, 6.2.0 from Play), who read the source and
arrived at the mechanism himself. The analysis below is his, verified against the tree; the scope
and the implementation notes are mine.

Symptom. In realtime mode the field cannot be corrected while dictation runs. Delete a word the
model got wrong — noise in the room, someone else talking — and the next update puts it back. Worse
at length: select all, delete, and the next update pastes the entire conversation so far back into
the field. The user ends up with a field full of text they had already removed, and no way to fix
anything without stopping first.

Mechanism. DictateController keeps realtimeShown, its own copy of what it believes is in the
field, and every write is computed against that copy rather than against the field:
setDictationPreview(full, realtimeShown), commitDictationFinal(outputText, realtimeShown),
clearDictationPreview(realtimeShown).

All three land in ImeDictationSink and none of them reads the field first:

  • applyDictationDiff (DictationSink.kt:143) keeps the common prefix with the remembered string
    and calls replaceTextBeforeCursor(old.length - cp, …) for the divergent tail.
  • commitDictationFinal (DictationSink.kt:126) does the same against prevText.
  • clearDictationPreview (DictationSink.kt:134) deletes prevText.length characters before the
    cursor outright.

So the moment the user edits by hand, the remembered string stops describing reality. The delete
count is then measured against text that is no longer there — it eats the user's own words — and the
tail that "diverges" is the whole transcript, which gets committed again.

The fix has a precedent in the same file. deleteLastText (DictationSink.kt:112) already
refuses to act when the field does not match: it only undoes when textBeforeSelection ends with
exactly what was inserted. The three realtime paths should make the same comparison and, when it
fails, accept the field as the new baseline instead of trying to reconcile against a copy that
is already wrong. Deleted text then stays deleted and the next words are appended at the cursor,
which is what the user expects.

The trap in doing that. EditorContent.textBeforeSelection is capped at 256 characters
(NumCharsBeforeCursor, AbstractEditorInstance.kt:56). realtimeShown routinely exceeds that in a
long dictation, so a plain endsWith(realtimeShown) would fail on length alone and re-baseline on
every single update — turning a correctness fix into a permanent one. The comparison has to be
bounded to the available window: compare the last n characters, where n is the smaller of the
window and the remembered string, and treat "the window is shorter than what we remember" as a match
rather than a mismatch.

Scope. ImeDictationSink only. AccessibilitySink ignores prevText and manages its own
preview through the service, and RecognitionSink has no field access at all — neither can drift
this way.

Worth deciding while in here: whether a re-baseline should also reset the realtime provider's
notion of the session, or only ours. Appending at the cursor with a provider that still holds the
full transcript means the next revision it sends can reach back across the user's edit again. My
inclination is that a re-baseline truncates what we will ever revise — anything before the edit
point is the user's now, not ours — but that needs a device test against a provider that revises
aggressively, not an argument.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions