Skip to content

The sort key of a node is spelt once: SimplifyHard -35% allocation - #1210

Merged
Rafael-SOWNet merged 1 commit into
masterfrom
speed-simplify
Sep 8, 2026
Merged

Rafael-SOWNet merged 1 commit into
masterfrom
speed-simplify

Conversation

@Rafael-SOWNet

Copy link
Copy Markdown
Member

Part of #746 — the standing condition on speed. Same method as #1205, #1207 and #1209: a temporary hook attributing allocated bytes, this time to the registration steps of Simplificator.Alternate on SimplifyHard.

What the attribution showed

The CanonicalOrder sort inside registration was 174 MB of the run's 735, about 2.6 MB per sort of a tree a few hundred nodes wide. Entity.SortHash(level) builds a node's key out of the keys of every node below it, and nothing kept them: one sort re-spelt each subtree once per ancestor, and the simplifier sorts what it registers at every pass of every level — in its own run and in the three nested runs its expand/factor/multiple-angle candidates cost.

What changes

The key is kept on the instance: one reference per node, an array made on the first sort, with the node the array was made for in its last slot, behind a struct that is equal to every other struct of its kind. Four shapes were measured and the three rejected ones are written beside the one kept:

shape SimplifyHard why not
three LazyPropertyA<string> slots 449 MB, 266 ms every node +~50 B: Derivate +4.7%, SolveMediumHard +5.6%, three more rows over the gate's 3%
a [ThreadStatic] dictionary for the run, keyed by reference 439 MB, 369 ms a lookup per node per sort costs more time than the strings it saves
a bare object?[]? field with an owner slot 30 GB, 57 s Entity is a record; the field took part in equality, a sorted tree stopped equalling its unsorted twin, and the simplifier never recognised a repeated candidate. Its unit test passed.
the same array behind an equality-neutral struct 443 MB, 258 ms kept

The owner slot matters on its own: WithCodomain is this with { Codomain = … } and a re-differentiation is this with { Iterations = … }, both of which copy every field of the record, while a number's key spells its codomain. A copy finds an array that is not its own and starts one.

Measured

Gate on one machine, both columns in one session:

baseline (574555bf) this change allocation time
SimplifyHard 680,457,400 B, 339 ms 442,898,152 B, 258 ms −34.9% 0.76x
SimplifyEasy 110,051 B, 66.4 µs 106,643 B, 68.8 µs −3.1% 1.04x
every solve entry, Derivate, ParseEasy +0.4% to +0.9% 1.00x

The fraction of a percent on the other rows is the one reference per node. Since the 1930th column: SimplifyHard −88%. The baseline is updated in the same change.

Answers

None move: the key is the same string, spelt once. SortKeyCacheTest pins that the key is the same on every ask and for a separately built equal tree, that a with copy with another codomain spells its own key, that sorting leaves equality and hash codes untouched, and that every thread gets the same key.

Checks

Full suite in two chunks: 9,666 passed, 0 failed (Calculus 1,322, the rest 8,344).

🤖 Generated with Claude Code

https://claude.ai/code/session_012sonx8iAspMiwRwokT1Ura

Attributing SimplifyHard's 735 MB to the registration steps of
Simplificator.Alternate found the CanonicalOrder sort at 174 MB, some
2.6 MB per sort of a tree a few hundred nodes wide. Entity.SortHash builds
a node's key out of the keys of every node below it, and nothing kept
them: one sort re-spelt each subtree once per ancestor, and the
simplifier sorts what it registers at every pass of every level, in its
own run and in the three nested runs its candidates cost.

The key is kept on the instance: one reference per node, an array made on
the first sort, with the node the array was made for in its last slot,
behind a struct that is equal to every other. Four measured shapes,
written beside it. A lazy slot per level cost every node some fifty
bytes, +4-6% on every solve and derivative entry of the gate. A dictionary
kept for the run, keyed by reference, saved the same bytes and cost a
lookup per node per sort, 369 ms against 266. A bare array field took
part in the record's equality, so a sorted tree stopped equalling its
unsorted twin and the simplifier never recognised a repeated candidate:
30 GB and 57 s. And a `with` copy (WithCodomain, a re-differentiation)
carries the original's fields while a number's key spells its codomain,
so the array names its owner and a copy starts one of its own.

Measured by the gate on one machine: SimplifyHard 680,457,400 ->
442,898,152 B and 339 -> 258 ms, SimplifyEasy -3.1%, every other entry
within +0.9%. The baseline moves with it. No answer moves: the key is the
same string, spelt once, and equality is untouched -- SortKeyCacheTest
pins both.

Part of #746.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012sonx8iAspMiwRwokT1Ura
@Rafael-SOWNet
Rafael-SOWNet merged commit ab00d7d into master Sep 8, 2026
31 checks passed
@Rafael-SOWNet
Rafael-SOWNet deleted the speed-simplify branch September 8, 2026 01:03
Rafael-SOWNet added a commit that referenced this pull request Sep 8, 2026
…was offered at: SimplifyHard -25% allocation (#1211)

Simplificator.Alternate offers three candidates after its level loop --
the expansion, the rule-based factorisation and the opened multiple
angles -- and re-simplifies each in full before the metric is asked,
because their cancellations only show once they are simplified. Each
re-simplification is a nested Alternate at the caller's own level, so at
level 4 the three cost three nested four-level searches: 566 MB of
SimplifyHard's 735 before the sort-key cache, 77% of the run, for
candidates the metric then ranked. Measured level by level with a
temporary hook, the third and fourth levels of a nested run register one
or two new trees where the first two register dozens, and cost the same
passes.

The candidates are re-simplified at level 2 now, the default, whatever
level the outer run was asked for; the outer loop keeps its level. A
caller at the default level sees no difference at all. A caller above it
gets a candidate simplified the way every default call simplifies, ranked
by the same metric against an outer run that still ran at the level asked
for.

Measured by the gate on one machine, on top of #1210: SimplifyHard
442,898,152 -> 331,683,360 B and 258 -> 190 ms, SimplifyEasy 106,643 ->
79,282 B and 68.8 -> 47.4 us; every other entry identical to the byte.
The baseline moves with it. No pinned answer moves: the callers above the
default level are the solver's alternative spellings, polynomial long
division's coefficients and whoever passes a level by hand, and the suite
covers the first two.

Part of #746.


Claude-Session: https://claude.ai/code/session_012sonx8iAspMiwRwokT1Ura

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant