Repository navigation
Two rule sets run on the matcher rather than on their switch - #1052
Merged
Merged
Conversation
#746 tier 1 records pattern matching as data (#248) as delivered and then says what is wrong with that: the engine exists and **nothing runs on it** -- every type in `Core/Transformations/Matching` is internal, five sets are expressed as data, and 0 of the 30 registered sets used the matcher at run time. The reason was a number. Swapping `DivisionPreparing` for its data form had been measured at about 5% of `Simplify`, and 5% for one of thirty sets is not affordable. #1050 changed that number by settling a pattern's determinism once instead of on every attempt. Re-measured against `Simplify` itself -- both arms in one process, plus a third arm that is the switch again as a control, because comparing two runs on this machine measures the machine: switch median 539 ms data median 536 ms -0.6% control median 533 ms -1.1% <- the switch against itself The change is smaller than this machine's disagreement with itself. Allocation is 38,181,736 B against 38,196,896 B, +0.04%. So `DivisionPreparing` and `CollapseMultipleFractions` now execute as data. They are the two whose data form is proven to agree with the switch it mirrors over generated expressions, with a minimum firing count so agreement cannot be vacuous -- which is the precondition for running one in place of the other. The switch stays, and is not dead code. `RewriteRule` carries a `PatternSource` and a `SourceLine` that `RuleRegistryGenerator` reads off a switch's arms, and it cannot yet read a rule written as data, so the addressable rules of these two sets still come from `Patterns` and describe arms the agreement test holds the matcher to. Teaching the generator to read `MatchedRules` is what would let the switch go, and is #825. Measured, not assumed: suite 8553/0, and rulecheck, canoncheck, casbench and simpsweep are byte-identical to master's committed reports -- rulecheck reporting 0 value changes across 1005 applications of 30 sets, which is the harness that attributes a change to a named set. Part of #746 tier 1 (#248). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjumi5K7fg8yx6UK1mZTQd
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of #746 tier 1 (#248). Depends on #1050, which is merged.
The gap this closes
#746's tier 1 row records pattern matching as data as delivered, and then says what is wrong with
that:
The blocker was that number, and it was a fair one: 5% for one of thirty sets does not scale to
thirty. #1050 changed it, by settling a pattern's determinism once rather than on every attempt.
Re-measured against
Simplify, not against a microbenchmarkA microbenchmark ratio is how the 5% became a surprise the first time, so this is measured against
Simplifyitself — both arms in one process, interleaved, with a third arm that is theswitchagain as a control, because comparing two runs on this machine measures the machine:
The control differs from its own source code by more than the change does. Allocation is
38,181,736 B against 38,196,896 B, +0.04%. Both arms were checked to give the same answer on
every input before either was timed.
What is exchanged, and why only these two
DivisionPreparingandCollapseMultipleFractionsare the two sets whose data form is proven toagree with the
switchit mirrors, over generated expressions, with a minimum firing count so theagreement cannot be vacuous —
MatchedRulesAgreeWithTheSwitchTest. That proof is the preconditionfor running one in place of the other, and the other three data sets do not have it:
PowerOfPoweris checked against a
switchthat does more than it, andPythagoreanIdentityhas nothing to agreewith by design.
The
switchstays, and is not dead codeRewriteRulecarries aPatternSourceand aSourceLine, whichRuleRegistryGeneratorreads offthe arms of a
switch. It has no way yet to read a rule written as data, so the addressable rules ofthese two sets still come from
Patterns— and the arms they describe are exactly the ones theagreement test holds the matcher to, rather than code nothing runs. Teaching the generator to read
MatchedRulesis what would let theswitchgo, and is #825.This is stated in the doc comment rather than left for a reader to discover, because "the registry
describes a
switchthat no longer executes" is the kind of thing that is true and invisible.Evidence
rulecheck,canoncheck,casbenchandsimpsweepre-run against this branch arebyte-identical to master's committed reports, ignoring the commit line and timings.
rulecheckreports 0 value changes across 1005 applications of 30 sets, and it is the harnessthat attributes a change to a named set — so a difference in either of these two would have been
named rather than averaged away.
After this, #746 tier 1's "0 of the 30 registered sets use the matcher at runtime" reads 2 of 30.