Skip to content

Send only valid UTF-8 so one bad string can't break queries on a source - #56

Merged
PetrHeinz merged 4 commits into
mainfrom
claude/scrub-invalid-utf8
Oct 2, 2026
Merged

PetrHeinz merged 4 commits into
mainfrom
claude/scrub-invalid-utf8

Conversation

@PetrHeinz

@PetrHeinz PetrHeinz commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

The HTTP device only relabels strings as UTF-8 (force_encoding) without checking them, and it never looks into Arrays or Hash keys. LogEntry also cuts long messages at byte 8,192, which can split a multibyte character in two. Better Stack stores a line with invalid UTF-8 as invalid JSON: its fields read empty, and every json.* query on the whole source then fails with ClickHouse Code 117. In the red-team, a single request with raw bytes in its path, logged by logtail-rack, was enough to break queries on a source.

  • Before encoding, every string in the line is made valid UTF-8: the message, nested fields, the context, Array elements and Hash keys.
    • UTF-8, binary (ASCII-8BIT) and US-ASCII strings are read as UTF-8, and invalid bytes become U+FFFD (scrub).
    • Strings in other encodings are converted with encode, so ISO-8859-1 "Montr\xE9al" arrives as Montréal. Characters with no UTF-8 equivalent become U+FFFD.
    • Ruby has no converter for 9 rare encodings, such as Windows-1258 and UTF-7. Strings in those are read as UTF-8 and scrubbed instead of raising.
  • LogEntry cuts a message longer than 8,192 bytes on a character boundary, dropping the bytes of the character the cut splits.

Behaviour and compatibility:

  • Valid UTF-8 strings are sent byte for byte as before.
  • Strings with invalid bytes now arrive as readable text with U+FFFD. Binary strings inside Arrays used to go out as msgpack bin values; they're now text like everywhere else.
  • Strings in other encodings now arrive converted instead of mislabelled.
  • In a message longer than 8,192 bytes, other invalid bytes are dropped by the same scrub instead of replaced. A long message labelled as binary is still cut at byte 8,192, and a character split there becomes U+FFFD, so it can end up 1 or 2 bytes over the limit.
  • logtail-rack needs no change: the strings it logs pass through the same code.

Follow-up after the E2E run of the release candidate (two more commits, the first one only tests and red on CI): force_utf8_encoding duplicated and scrubbed every string, also the ones that are valid UTF-8 already, which are nearly all, and together with #54 encoding a log line got several times slower than on 0.1.20. It now returns a string that is valid UTF-8, or US-ASCII without bytes above 127, as it is, and copies only the others. msgpack sends both the same way, so nothing changes in what arrives. Encoding a typical line, incl. deflate, on Ruby 3.4.9 with 1,000 log lines from the released logtail-rack 0.2.8 (0.1.20: 3.42 µs):

Targets the logtail 0.1.21 patch release. No dependencies on the other open PRs. It pairs with #54, which converts the values msgpack can't encode in the same code path; the two merge cleanly in either order and pass the suite together.

The first commit only adds the tests and is expected to fail on CI; the fix follows in the next commit.

🤖 Generated with Claude Code

PetrHeinz and others added 2 commits October 1, 2026 18:46
The HTTP device only relabels strings as UTF-8 without replacing
invalid bytes, skips Array elements and Hash keys, and LogEntry cuts
long messages in the middle of a multibyte character. Better Stack
then stores the line as invalid JSON. These tests pin the expected
behaviour: invalid bytes become U+FFFD everywhere in the line, strings
in other encodings are converted, and a long message is cut on a
character boundary.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
force_utf8_encoding only relabelled strings as UTF-8, without replacing
invalid bytes, and skipped Array elements and Hash keys. LogEntry cut
long messages with byteslice, which can split a multibyte character.
Better Stack stores such a line as invalid JSON, and every json.* query
on the source then fails.

- UTF-8, binary and US-ASCII strings are read as UTF-8 and scrubbed:
  invalid bytes become U+FFFD.
- Strings in other encodings are converted with encode, replacing what
  has no UTF-8 equivalent. The few encodings Ruby has no converter for
  are read as UTF-8 and scrubbed instead of raising.
- Hash keys and Array elements are handled recursively, like values.
- A message longer than 8,192 bytes is cut on a character boundary.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
PetrHeinz and others added 2 commits October 2, 2026 08:10
force_utf8_encoding duplicates and scrubs every string, also the ones that
are valid UTF-8 already, which are nearly all. That made encoding a log
line several times slower than on 0.1.20.

The second test pins that a US-ASCII string with bytes above 127 is still
treated as UTF-8.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
force_utf8_encoding returns a string that is valid UTF-8, or US-ASCII
without bytes above 127, as it is. Only the others are copied and scrubbed
or converted, as before. msgpack sends both the same way, so nothing
changes in what arrives.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@PetrHeinz
PetrHeinz marked this pull request as ready for review October 2, 2026 09:08
@PetrHeinz
PetrHeinz merged commit bb8f09f into main Oct 2, 2026
13 checks passed
@PetrHeinz
PetrHeinz deleted the claude/scrub-invalid-utf8 branch October 2, 2026 11:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant