Send only valid UTF-8 so one bad string can't break queries on a source - #56
Merged
Merged
Conversation
The HTTP device only relabels strings as UTF-8 without replacing invalid bytes, skips Array elements and Hash keys, and LogEntry cuts long messages in the middle of a multibyte character. Better Stack then stores the line as invalid JSON. These tests pin the expected behaviour: invalid bytes become U+FFFD everywhere in the line, strings in other encodings are converted, and a long message is cut on a character boundary. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
force_utf8_encoding only relabelled strings as UTF-8, without replacing invalid bytes, and skipped Array elements and Hash keys. LogEntry cut long messages with byteslice, which can split a multibyte character. Better Stack stores such a line as invalid JSON, and every json.* query on the source then fails. - UTF-8, binary and US-ASCII strings are read as UTF-8 and scrubbed: invalid bytes become U+FFFD. - Strings in other encodings are converted with encode, replacing what has no UTF-8 equivalent. The few encodings Ruby has no converter for are read as UTF-8 and scrubbed instead of raising. - Hash keys and Array elements are handled recursively, like values. - A message longer than 8,192 bytes is cut on a character boundary. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
force_utf8_encoding duplicates and scrubs every string, also the ones that are valid UTF-8 already, which are nearly all. That made encoding a log line several times slower than on 0.1.20. The second test pins that a US-ASCII string with bytes above 127 is still treated as UTF-8. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
force_utf8_encoding returns a string that is valid UTF-8, or US-ASCII without bytes above 127, as it is. Only the others are copied and scrubbed or converted, as before. msgpack sends both the same way, so nothing changes in what arrives. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The HTTP device only relabels strings as UTF-8 (
force_encoding) without checking them, and it never looks into Arrays or Hash keys.LogEntryalso cuts long messages at byte 8,192, which can split a multibyte character in two. Better Stack stores a line with invalid UTF-8 as invalid JSON: its fields read empty, and everyjson.*query on the whole source then fails with ClickHouse Code 117. In the red-team, a single request with raw bytes in its path, logged by logtail-rack, was enough to break queries on a source.scrub).encode, so ISO-8859-1"Montr\xE9al"arrives asMontréal. Characters with no UTF-8 equivalent become U+FFFD.LogEntrycuts a message longer than 8,192 bytes on a character boundary, dropping the bytes of the character the cut splits.Behaviour and compatibility:
binvalues; they're now text like everywhere else.scrubinstead of replaced. A long message labelled as binary is still cut at byte 8,192, and a character split there becomes U+FFFD, so it can end up 1 or 2 bytes over the limit.Follow-up after the E2E run of the release candidate (two more commits, the first one only tests and red on CI):
force_utf8_encodingduplicated and scrubbed every string, also the ones that are valid UTF-8 already, which are nearly all, and together with #54 encoding a log line got several times slower than on 0.1.20. It now returns a string that is valid UTF-8, or US-ASCII without bytes above 127, as it is, and copies only the others. msgpack sends both the same way, so nothing changes in what arrives. Encoding a typical line, incl. deflate, on Ruby 3.4.9 with 1,000 log lines from the released logtail-rack 0.2.8 (0.1.20: 3.42 µs):force_utf8_encoding;Targets the logtail 0.1.21 patch release. No dependencies on the other open PRs. It pairs with #54, which converts the values msgpack can't encode in the same code path; the two merge cleanly in either order and pass the suite together.
The first commit only adds the tests and is expected to fail on CI; the fix follows in the next commit.
🤖 Generated with Claude Code