Skip to content

Publish the comparisons against other systems (#184) - #1030

Merged
Rafael-SOWNet merged 5 commits into
masterfrom
docs/publish-the-comparisons
Aug 23, 2026
Merged

Rafael-SOWNet merged 5 commits into
masterfrom
docs/publish-the-comparisons

Conversation

@Rafael-SOWNet

Copy link
Copy Markdown
Member

Closes nothing on its own; it is item 40 of #746 -- "compare us against another CAS on a fixed corpus and publish the table (#184)" -- whose remaining half was the publishing. The measurements existed in the analysis workspace and nothing in this repository said what they found.

New file Sources/AngouriMath/Docs/Usage/Comparison.md, linked from the root README.md under How does it compare?. No code changes, so no BREAKING-CHANGES.md entry.

What it publishes

comparison the other side this library measured
capability, output and speed Math.NET Symbolics 0.25.0, Symbolism 1.0.4, on .NET 10.0.10 6b93b40, the 2.3.0 commit 2026-08-23
80 feature probes over 23 areas SymPy 1.14.0 6b93b40 2026-08-23
1,774 integration problems Rubi's test-suite, twelve files 3ac24bc 2026-08-23

3ac24bc is three commits before the tag and the three differ only in BREAKING-CHANGES.md, the performance table and the package version, so no library code separates them.

The losing rows are in it

  • Math.NET is faster on all four operations both libraries can do, by 5.2x to 110.4x.
  • Math.NET's answer is shorter on 4 of the 9 shared tasks where the two differ.
  • 35 of 80 SymPy probes find no public member at all here; 5 more are implemented but internal.
  • 604 of 1,774 Rubi problems answered, 34.0%.
  • The single errors row is x^2 - 4 > 0, where this library raises where SymPy returns two intervals.

Every sentence of judgement is a count taken from the table directly above it, because this repository has had the other kind: a generated report once emitted a hardcoded sentence naming two outputs as evidence, under a table that contradicted it.

Two claims from the source reports are deliberately not repeated. comparison.md's prose says the d/dx ln(x)/x row records x > 0; the table shows provided not x = 0, and the table wins. Its nuget download counts are undated and measure adoption, not the library.

What it says it does not establish

answers in the SymPy table is computed from this library's answer alone, so it is not agrees, and no disagreement on the page is adjudicated. absent is a keyword search of the public surface, so it is a lower bound. 604/1774 is a statement about Rubi's textbook problems at a 5-second budget -- set against the 116/119 the library's own corpus scores the same day, the pair says more about who wrote each list than about integration.

And the reproduction section says plainly that the harnesses are not in this repository, not in the solution and not in CI, so dotnet test reproduces no figure on the page; what does run on every commit is Sources/Tests/UnitTests/Corpus, which is a gate and not the source of anything here.

Note for whoever owns AGENTS.md

Its Where things are written down table wants a row for this file. Not added here -- that file is owned by another change in flight.

🤖 Generated with Claude Code

https://claude.ai/code/session_01Bjumi5K7fg8yx6UK1mZTQd

darkfader and others added 3 commits August 23, 2026 12:33
The measurements existed and nothing in the repository said what they found.
Docs/Usage/Comparison.md reproduces three of them -- the head-to-head against
Math.NET Symbolics 0.25.0 and Symbolism 1.0.4, the 80 feature probes against
SymPy 1.14.0, and 1,774 problems of Rubi's integration suite -- with the
version, date and library commit each was taken at, and with the rows this
library loses reproduced alongside the rows it wins.

Every sentence of judgement on the page is a count taken from the table
directly above it: 3 of 12 shared tasks identical ignoring spacing, Math.NET
shorter on 4 of the remaining 9 and all four of those rows ones where a
`provided` clause is attached here; eleven of 24 capability rows present here
and in neither of the other two, and two the other way; Math.NET faster on
all four operations both libraries can do.

Three caveats the tables do not state and a reader needs. SymPy's verdict
column is computed from this library's answer alone, so `answers` is not
`agrees` and no disagreement on that page is adjudicated. `absent` is a
keyword search of the public surface, which makes it a lower bound. And
604/1774 is a statement about Rubi's textbook problems at a 5s budget -- held
against the 116/119 the library's own corpus scores, the pair says more about
who wrote each list than about integration.

The harnesses are outside this repository and are not in CI, so the page says
that plainly rather than implying `dotnet test` reproduces a figure on it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjumi5K7fg8yx6UK1mZTQd
The SymPy table said internal 5 and unclear 2; work/sympyparity.md says 3 and
4. Two probes were attributed to the wrong verdict, and nothing caught it
because both readings sum to the same 80 rows.

Re-checked every other figure on the page against its source in the same pass:
3 of 12 shared tasks identical, Math.NET shorter on 4, and Rubi 604 of 1,774
at 34.0% all match.
@Rafael-SOWNet

Copy link
Copy Markdown
Member Author

Two figures on this page were wrong, and are now corrected in 472e8da2/the commit above.

The SymPy verdict table said internal 5 and unclear 2. work/sympyparity.md — the report the
page is transcribed from — says internal 3 and unclear 4. Two probes were attributed to
the wrong verdict.

Worth recording why it survived review: both readings sum to 80, which is the row count stated
in the sentence above the table, so the obvious sanity check passes on the wrong numbers. The error
was only visible by reading the source report, which is what regenerating the harness at
e4eefde0 prompted.

Every other figure on the page has now been checked against its source in the same pass and matches:
3 of 12 shared tasks identical ignoring spacing, Math.NET shorter on 4 of the remaining 9, and Rubi
604 of 1,774 at 34.0%.

One still open: the Rubi figure is being re-measured right now at --budget=5 against merged master.
If it moves, this page moves with it before merge.

Also merged master in, so this is up to date with all nine of #1009–#1017.

The page carried 604 of 1774 at 3ac24bc, three commits before the 2.3.0 tag.
The suite is downloaded rather than vendored, so that harness runs on its own
schedule and had not been re-run since. It has now, at e4eefde -- master with
nine pull requests merged -- and answers 602 of 1774.

The two-problem difference is not a change in the integrator and the page now
says so instead of leaving a reader to infer a regression. All three cases that
moved between the runs moved into the timeout bucket, none became unevaluated
and none became wrong, and timed alone against this same build they take
1744 ms, 1579 ms and 835 ms -- none near the 5 second budget. intbench cannot
abort a thread, so a case that does time out leaks one and the leaked threads
slow every problem after it. The rate is therefore +/-3 problems and the page
says to read the wrong-answer count, which is 0, as the number that means
something.

Every per-file row re-taken from the report rather than edited by hand.
@Rafael-SOWNet

Copy link
Copy Markdown
Member Author

The Rubi row is now measured on merged master, and the page says what its rate is worth.

It carried 604 of 1,774 taken at 3ac24bc2, three commits before the 2.3.0 tag. That harness runs on
its own schedule because the suite is downloaded rather than vendored, and it had not been re-run
since. It has now, at e4eefde0 — master with all nine of #1009–#1017 merged — and answers
602 of 1,774 (33.9%), 0 wrong, 0 errors, 46 timeouts.

The two-problem difference is not the integrator, and the page now says so

All three cases that moved between the runs moved into the timeout bucket. None became
unevaluated, none became wrong. Timed on their own against the same build:

problem alone harness verdict
Timofeev Problems:660 — x^2*sin(x)^6 1,744 ms was solved, now timeout
Timofeev Problems:278 1,579 ms was solved, now timeout
Welz Problems:13 835 ms was unsolved, now timeout

None is near the 5-second budget. The mechanism is in intbench's own note at Program.cs:102:
.NET cannot abort a thread, so a case that times out leaks one, and the leaked threads slow every
problem after it. The page now tells a reader to hold the rate to ±3 problems and to read the
wrong-answer count — 0 — as the figure that means something.

That is #746 item 33's argument in one number: a wall-clock verdict is time-dependent by
construction, which is why the budget wants to be in work units (#1035).

Also in this push

Every per-file row was re-taken from work/intcoverage.md rather than hand-edited, and I checked all
twelve programmatically against the report — no mismatches. Combined with the SymPy correction
earlier, every figure on this page has now been verified against its source.

… I edited by hand

A third run of the same commit answered 604 with 44 timeouts, against 602 and
46 from the run before it. Restoring the page to 604 I put the Welz row back to
its old 97/4 from memory of the earlier column -- and the report says 96/5,
because Welz 13 timed out in both runs and only the two Timofeev problems came
back. Every per-file row is now substituted from work/intcoverage.md rather
than edited, and the twelve rows sum to the headline in all seven columns.

The variance paragraph is stronger for it: two runs of one commit, same slice
and same budget, differing by two problems, is a better argument than two runs
of two commits. 0 wrong in every run.
@Rafael-SOWNet
Rafael-SOWNet merged commit f995d72 into master Aug 23, 2026
31 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants