Skip to content

Every registered rule set runs as data (#825) - #1192

Merged
Rafael-SOWNet merged 1 commit into
masterfrom
tier2-bookkeeping
Sep 6, 2026
Merged

Rafael-SOWNet merged 1 commit into
masterfrom
tier2-bookkeeping

Conversation

@Rafael-SOWNet

Copy link
Copy Markdown
Member

The three CanonicalOrder sets were the last to run their switch. MatchedRules.Sort(level) had
existed as their data form, held to Patterns.SortRules(level) by MatchedRulesAgreeWithTheSwitchTest
at a hundred firings per level — and been run by nothing. They are repointed at it.

30 of 30 registered sets now execute the matcher, and 0 take their Rules from
RuleRegistryGenerator.
Both were 27 and 29 when RuleMetadataTest's remark was written; it says
so now.

Measured, not inferred from the agreement test

8,310 generated inputs simplified on a build of master and on this branch, output diffed: zero
differences.
The agreement test says the two spellings agree rule-by-rule; this says Simplify
agrees at the entry point, which is the thing a caller sees.

And it is not free

The write-up of the last rule-set exchange in this repository ends with "a run of individually-free
steps needs one measurement against where it started, or it has not been measured at all."
So,
allocation — master (eb3a0447) against this change, one benchmark on one machine:

benchmark master this
SimplifyHard 3,655,695,192 3,662,375,184 +0.18%
SolveHard 1,450,058,856 1,452,704,384 +0.18%
SolveMediumHard 165,055,344 165,611,184 +0.34%
SimplifyEasy 128,099 128,460 +0.28%
CompileHard 20,753 20,434 −1.5% — inside that family's ±2% floor, means nothing

Real, because those Simplify/Solve rows reproduce to one part in 10⁵. Inside PerformanceGate's
3%. And the third small step since 2.4.0 on this family, after #1176 and #1177 — roughly half a
per cent cumulative. The Sort set runs on every node of every pass, so a matcher dispatch where a
switch used to be is the likeliest source; it is the price of the set being addressable, named and
tiered like the other twenty-seven. Stated so it is a known cost rather than a surprise later.

(The A/B's first table came out empty because I restored key-commits.txt while the script was still
running — it reads the list a second time at combine time. The measurements were on disk; only the
table had to be redrawn. Worth knowing before the next person restores that file mid-run.)

The gate tripped, correctly

RuleMetadataTest.EveryRuleOfARepointedSetCarriesItsIdentity requires that a set is repointed only
once every rule of it carries an identity, so the registry gains descriptions rather than trading
them.
The nine Sort rules had none — they rendered as Xorf x => (built by code). I had repointed
a set that wasn't ready. The fix is at the source: each of the nine carries a description now.
The gate is untouched.

Five pins moved because the exchange succeeded

Each to the number its failure reported:

pin was is
sets describing what they run 27 30
addressable rules 315 321
rules carrying an identity (two tests, and WritingARule.md) 294 321

One pin was wrong while passing. "How many registered sets run the matcher" matched
MatchedRules names — and CanonicalOrder runs under the name Sort, so it said 27 while the truth
was 30, for exactly as long as that family went unlisted. It has the same special case
CommonDenominator already had.

StepAsASentenceTest asserted that some set was still generator-described and that its names did
not pass for prose. That population is now empty, and the assertion flips to say so rather than
being deleted: a set falling back to the generator is a regression this should catch.

On the "six orphan sets"

The audit that prompted this named six MatchedRules sets as orphans. Measured: three are the
CommonDenominator family, wired all along. The other three — PowerOfPower, PythagoreanIdentity,
SharedFactor — are single-rule fixtures that ReversibleRuleTest, GatheredMatchingTest and the
e-match tests use as subjects; deleting them deletes the evidence. Two duplicate registry rules under
near-identical names. SharedFactor's a-shared-factor-comes-out-of-a-sum appears nowhere else,
which is worth its own look and is not taken here.

Part of #825 and #746.

Checks

Full suite 9571 passed, 0 failed, 14 skipped. 8,310-input sweep: 0 changed answers.

🤖 Generated with Claude Code

https://claude.ai/code/session_012sonx8iAspMiwRwokT1Ura

The three CanonicalOrder sets were the last to run their switch. MatchedRules.Sort(level)
had existed as their data form, held to Patterns.SortRules(level) by
MatchedRulesAgreeWithTheSwitchTest at a hundred firings per level, and been run by
nothing. They are repointed at it, through the constructor that takes a MatchedRuleSet
and so runs and describes the same object. 30 of 30 registered sets now execute the
matcher and 0 take their Rules from RuleRegistryGenerator; both numbers were 27 and 29
when RuleMetadataTest's remark was written, and it says so now.

Correctness at the entry point was measured rather than inferred from the agreement
test: 8310 generated inputs simplified on a build of master and on this one, diffed,
zero differences.

And it is not free, which is the thing the last exchange's write-up says to state.
Allocation, master against this commit, one benchmark on one machine:

    SimplifyHard      3,655,695,192 -> 3,662,375,184   +0.18%
    SolveHard         1,450,058,856 -> 1,452,704,384   +0.18%
    SolveMediumHard     165,055,344 ->   165,611,184   +0.34%
    SimplifyEasy            128,099 ->       128,460   +0.28%

Real -- those rows reproduce to a part in a hundred thousand -- and inside
PerformanceGate's 3%. The compile rows moved -1.5%, which is inside their own 2% floor
and means nothing. The Sort set runs on every node of every pass, so a matcher dispatch
where a switch used to be is the likeliest source; it is the price of the set being
addressable, named and tiered like the other twenty-seven, and since 2.4.0 this family
has now paid roughly half a per cent across three such steps.

The repoint tripped the gate it should have. RuleMetadataTest requires that a set is
repointed only once every rule of it carries an identity, so the registry gains
descriptions rather than trading them -- and the nine Sort rules had none, rendering as
`Xorf x => (built by code)`. They carry one each now. The fix is at the source, not the
gate.

Five census pins moved because the exchange succeeded, each to the number the failure
reported: 27 sets describing what they run is 30; 315 addressable rules is 321; 294
rules carrying an identity is 321, in two tests and in WritingARule.md. One pin was
wrong while passing: "how many registered sets run the matcher" matched MatchedRules
names, and CanonicalOrder runs under the name Sort, so it said 27 while the truth was 30.
It has the same special case CommonDenominator already had.

StepAsASentenceTest asserted that some set was still generator-described and that its
names did not pass for prose. That population is empty now, and the assertion flips to
say so rather than being deleted: a set falling back to the generator is a regression
this should catch.

The audit that prompted this called six MatchedRules sets orphans. Measured: three of
those are the CommonDenominator family, wired all along; the other three --
PowerOfPower, PythagoreanIdentity, SharedFactor -- are single-rule fixtures that
ReversibleRuleTest, GatheredMatchingTest and the e-match tests use as subjects, and
deleting them deletes the evidence. Two duplicate registry rules under near-identical
names. SharedFactor's `a-shared-factor-comes-out-of-a-sum` appears nowhere else, which
is worth its own look and is not taken here.

Part of #825 and #746.

Full suite 9571 passed, 0 failed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012sonx8iAspMiwRwokT1Ura
@Rafael-SOWNet
Rafael-SOWNet merged commit e87cd5d into master Sep 6, 2026
31 checks passed
@Rafael-SOWNet
Rafael-SOWNet deleted the tier2-bookkeeping branch September 6, 2026 20:23
Rafael-SOWNet added a commit that referenced this pull request Sep 7, 2026
…hine (#1204)

The release checklist owes a measured performance column with the previous one
re-measured on the same machine. This is that column for 308384b: every release from
v2.1.0 to v2.4.0 and master, each measured on 2026-09-07 with nothing else on the
machine, one entry per run because a single command here is capped at ten minutes.

Since the 1915th, SimplifyHard is +0.18%, SolveHard +0.18% and SolveMediumHard
+0.34% -- to the digit the cost recorded for #1192 when it repointed the CanonicalOrder
sets at their data form -- so the thirteen pull requests after it read as
allocation-neutral on these benchmarks, and the section says that is a reading of two
agreeing measurements rather than a bisection. SimplifyEasy's +360 bytes is the one
figure that reading does not cover, and is left unattributed. Since 2.4.0, the pair a
release publishes, the Simplify and Solve family is up between 0.28% and 0.84%, the
compile rows sit inside their two-percent floor, and everything else is flat.

v2.4.0 measured twice a fortnight apart agrees to the byte on SimplifyHard and
SolveMediumHard and to 0.0005% on SolveHard, so the file's determinism claim holds
again with the same rider on the compile rows.

Part of #746, and item 3 of the release checklist.


Claude-Session: https://claude.ai/code/session_012sonx8iAspMiwRwokT1Ura

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant