Repository navigation
Every registered rule set runs as data (#825) - #1192
Merged
Merged
Conversation
The three CanonicalOrder sets were the last to run their switch. MatchedRules.Sort(level)
had existed as their data form, held to Patterns.SortRules(level) by
MatchedRulesAgreeWithTheSwitchTest at a hundred firings per level, and been run by
nothing. They are repointed at it, through the constructor that takes a MatchedRuleSet
and so runs and describes the same object. 30 of 30 registered sets now execute the
matcher and 0 take their Rules from RuleRegistryGenerator; both numbers were 27 and 29
when RuleMetadataTest's remark was written, and it says so now.
Correctness at the entry point was measured rather than inferred from the agreement
test: 8310 generated inputs simplified on a build of master and on this one, diffed,
zero differences.
And it is not free, which is the thing the last exchange's write-up says to state.
Allocation, master against this commit, one benchmark on one machine:
SimplifyHard 3,655,695,192 -> 3,662,375,184 +0.18%
SolveHard 1,450,058,856 -> 1,452,704,384 +0.18%
SolveMediumHard 165,055,344 -> 165,611,184 +0.34%
SimplifyEasy 128,099 -> 128,460 +0.28%
Real -- those rows reproduce to a part in a hundred thousand -- and inside
PerformanceGate's 3%. The compile rows moved -1.5%, which is inside their own 2% floor
and means nothing. The Sort set runs on every node of every pass, so a matcher dispatch
where a switch used to be is the likeliest source; it is the price of the set being
addressable, named and tiered like the other twenty-seven, and since 2.4.0 this family
has now paid roughly half a per cent across three such steps.
The repoint tripped the gate it should have. RuleMetadataTest requires that a set is
repointed only once every rule of it carries an identity, so the registry gains
descriptions rather than trading them -- and the nine Sort rules had none, rendering as
`Xorf x => (built by code)`. They carry one each now. The fix is at the source, not the
gate.
Five census pins moved because the exchange succeeded, each to the number the failure
reported: 27 sets describing what they run is 30; 315 addressable rules is 321; 294
rules carrying an identity is 321, in two tests and in WritingARule.md. One pin was
wrong while passing: "how many registered sets run the matcher" matched MatchedRules
names, and CanonicalOrder runs under the name Sort, so it said 27 while the truth was 30.
It has the same special case CommonDenominator already had.
StepAsASentenceTest asserted that some set was still generator-described and that its
names did not pass for prose. That population is empty now, and the assertion flips to
say so rather than being deleted: a set falling back to the generator is a regression
this should catch.
The audit that prompted this called six MatchedRules sets orphans. Measured: three of
those are the CommonDenominator family, wired all along; the other three --
PowerOfPower, PythagoreanIdentity, SharedFactor -- are single-rule fixtures that
ReversibleRuleTest, GatheredMatchingTest and the e-match tests use as subjects, and
deleting them deletes the evidence. Two duplicate registry rules under near-identical
names. SharedFactor's `a-shared-factor-comes-out-of-a-sum` appears nowhere else, which
is worth its own look and is not taken here.
Part of #825 and #746.
Full suite 9571 passed, 0 failed.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012sonx8iAspMiwRwokT1Ura
Rafael-SOWNet
added a commit
that referenced
this pull request
Sep 7, 2026
…hine (#1204) The release checklist owes a measured performance column with the previous one re-measured on the same machine. This is that column for 308384b: every release from v2.1.0 to v2.4.0 and master, each measured on 2026-09-07 with nothing else on the machine, one entry per run because a single command here is capped at ten minutes. Since the 1915th, SimplifyHard is +0.18%, SolveHard +0.18% and SolveMediumHard +0.34% -- to the digit the cost recorded for #1192 when it repointed the CanonicalOrder sets at their data form -- so the thirteen pull requests after it read as allocation-neutral on these benchmarks, and the section says that is a reading of two agreeing measurements rather than a bisection. SimplifyEasy's +360 bytes is the one figure that reading does not cover, and is left unattributed. Since 2.4.0, the pair a release publishes, the Simplify and Solve family is up between 0.28% and 0.84%, the compile rows sit inside their two-percent floor, and everything else is flat. v2.4.0 measured twice a fortnight apart agrees to the byte on SimplifyHard and SolveMediumHard and to 0.0005% on SolveHard, so the file's determinism claim holds again with the same rider on the compile rows. Part of #746, and item 3 of the release checklist. Claude-Session: https://claude.ai/code/session_012sonx8iAspMiwRwokT1Ura Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
This was referenced Sep 7, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The three
CanonicalOrdersets were the last to run theirswitch.MatchedRules.Sort(level)hadexisted as their data form, held to
Patterns.SortRules(level)byMatchedRulesAgreeWithTheSwitchTestat a hundred firings per level — and been run by nothing. They are repointed at it.
30 of 30 registered sets now execute the matcher, and 0 take their
RulesfromRuleRegistryGenerator. Both were 27 and 29 whenRuleMetadataTest's remark was written; it saysso now.
Measured, not inferred from the agreement test
8,310 generated inputs simplified on a build of
masterand on this branch, output diffed: zerodifferences. The agreement test says the two spellings agree rule-by-rule; this says
Simplifyagrees at the entry point, which is the thing a caller sees.
And it is not free
The write-up of the last rule-set exchange in this repository ends with "a run of individually-free
steps needs one measurement against where it started, or it has not been measured at all." So,
allocation —
master(eb3a0447) against this change, one benchmark on one machine:SimplifyHardSolveHardSolveMediumHardSimplifyEasyCompileHardReal, because those Simplify/Solve rows reproduce to one part in 10⁵. Inside
PerformanceGate's3%. And the third small step since 2.4.0 on this family, after #1176 and #1177 — roughly half a
per cent cumulative. The
Sortset runs on every node of every pass, so a matcher dispatch where aswitchused to be is the likeliest source; it is the price of the set being addressable, named andtiered like the other twenty-seven. Stated so it is a known cost rather than a surprise later.
(The A/B's first table came out empty because I restored
key-commits.txtwhile the script was stillrunning — it reads the list a second time at combine time. The measurements were on disk; only the
table had to be redrawn. Worth knowing before the next person restores that file mid-run.)
The gate tripped, correctly
RuleMetadataTest.EveryRuleOfARepointedSetCarriesItsIdentityrequires that a set is repointed onlyonce every rule of it carries an identity, so the registry gains descriptions rather than trading
them. The nine
Sortrules had none — they rendered asXorf x => (built by code). I had repointeda set that wasn't ready. The fix is at the source: each of the nine carries a description now.
The gate is untouched.
Five pins moved because the exchange succeeded
Each to the number its failure reported:
WritingARule.md)One pin was wrong while passing. "How many registered sets run the matcher" matched
MatchedRulesnames — andCanonicalOrderruns under the nameSort, so it said 27 while the truthwas 30, for exactly as long as that family went unlisted. It has the same special case
CommonDenominatoralready had.StepAsASentenceTestasserted that some set was still generator-described and that its names didnot pass for prose. That population is now empty, and the assertion flips to say so rather than
being deleted: a set falling back to the generator is a regression this should catch.
On the "six orphan sets"
The audit that prompted this named six
MatchedRulessets as orphans. Measured: three are theCommonDenominatorfamily, wired all along. The other three —PowerOfPower,PythagoreanIdentity,SharedFactor— are single-rule fixtures thatReversibleRuleTest,GatheredMatchingTestand thee-match tests use as subjects; deleting them deletes the evidence. Two duplicate registry rules under
near-identical names.
SharedFactor'sa-shared-factor-comes-out-of-a-sumappears nowhere else,which is worth its own look and is not taken here.
Part of #825 and #746.
Checks
Full suite 9571 passed, 0 failed, 14 skipped. 8,310-input sweep: 0 changed answers.
🤖 Generated with Claude Code
https://claude.ai/code/session_012sonx8iAspMiwRwokT1Ura