⚡ Bolt: optimize trie traversal and bloom filter hashing - #68
gregyjames wants to merge 2 commits into
Conversation
- Replace scalar multiplication inside Bloom filter loop with additive accumulation (`hash += h2`) - Store direct 0..25 bit indices in character normalization buffer to eliminate ASCII offset math during trie traversal - Compact Trie Node struct by encoding `end_of_word` into bit 31 of `children_mask` and removing unused struct fields Co-authored-by: google-labs-jules[bot] <161369871+google-labs-jules[bot]@users.noreply.github.com>
|
👋 Jules, reporting for duty! I'm here to lend a hand with this pull request. When you start a review, I'll add a 👀 emoji to each comment to let you know I've read it. I'll focus on feedback directed at me and will do my best to stay out of conversations between you and other bots or reviewers to keep the noise down. I'll push a commit with your requested changes shortly after. Please note there might be a delay between these steps, but rest assured I'm on the job! For more direct control, you can switch me to Reactive Mode. When this mode is on, I will only act on comments where you specifically mention me with New to Jules? Learn more at jules.google/docs. For security, I will only act on instructions from the user who triggered this task. |
|
Warning Review limit reachedNext included review available in 50 minutes. View limit detailsLimit details: You’ve used the included review currently available. This review ran on the open-source allowance, not this organization's plan, because the pull request author doesn't have an assigned seat. Waiting won't change this — ask an organization admin to assign them a seat, or add seats in Billing if every seat is already assigned, then retry. Review configuration: ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: 📒 Files selected for processing (1)
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: ⛔ Files ignored due to path filters (1)
📒 Files selected for processing (4)
Included review availability: Your plan provides up to 1 included review per hour; 0 remain after this review. 📝 WalkthroughWalkthroughThe trie now stores word-end state in its child bitmask and uses direct character indices during traversal. Bloom filter indexing uses incremental hashing. The repository also ignores generated benchmark and build directories. ChangesHyperTrie optimizations
Repository housekeeping
Priority: ⬇️ Low Estimated code review effort: 3 (Moderate) | ~20 minutes Change: Refactor 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
Full details: Docstring CoverageExplanation Docstring coverage is 25.00% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 12 functions across 2 files. (2 skipped: 2 unsupported.) ✨ Finishing Touches 💡 1📝 Generate docstrings 💡
🧪 Generate unit tests (beta)
Comment |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #68 +/- ##
==========================================
+ Coverage 95.83% 98.22% +2.39%
==========================================
Files 3 3
Lines 600 620 +20
Branches 600 620 +20
==========================================
+ Hits 575 609 +34
+ Misses 19 5 -14
Partials 6 6
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. |
- Replace scalar multiplication inside Bloom filter loop with additive accumulation (`hash += h2`) - Store direct 0..25 bit indices in character normalization buffer to eliminate ASCII offset math during trie traversal - Compact Trie Node struct by encoding `end_of_word` into bit 31 of `children_mask` and removing unused struct fields - Add unit tests for long strings (> 64 bytes) and non-alphabetic character filtering Co-authored-by: google-labs-jules[bot] <161369871+google-labs-jules[bot]@users.noreply.github.com>
Implemented three targeted optimizations in native Rust HyperTrie backend:
i * h2inside BloomFilter hashing loop with additive updatehash += h2.+ b'a'and- b'a'arithmetic operations during Trie traversal.Nodestruct from 112 bytes to 108 bytes by encodingend_of_wordinto bit 31 ofchildren_maskand removing unused fields.Impact: Reduces native HyperTrie benchmark execution time from 22.91 ms to 21.66 ms (~5.5% speedup). All unit tests pass.
PR created automatically by Jules for task 7413189709073019543 started by @gregyjames
Summary by CodeRabbit
Performance
Maintenance