Skip to content

The integrator stops answering the same question over and over (#1156) - #1157

Merged
Rafael-SOWNet merged 1 commit into
masterfrom
integral-answers-kept
Sep 4, 2026
Merged

Rafael-SOWNet merged 1 commit into
masterfrom
integral-answers-kept

Conversation

@Rafael-SOWNet

Copy link
Copy Markdown
Member

Simplify of an integral it cannot evaluate took a minute and a half. It now takes about a second.

master now
sin(x)/(x^2 + 1)^2 114 851 ms 1163 ms 99×
e^x/(x^2 + 1)^2 85 683 ms 1517 ms 56×
ln(x)/(x^2 + 1)^2 51 072 ms 188 ms 272×
1/(x^4 + 1)^2 3811 ms 9 ms 423×

A single cold Integrate, in a fresh process, is faster too — on all nine tried:

master now
sqrt(tan(x)) 154.5 ms 79.1 ms
x^2*sin(x) 21.0 ms 10.2 ms
1/(x^3 + 1) 49.8 ms 31.2 ms
x^2/(x^2 + 2)^2 17.4 ms 12.0 ms

The integrator was asked the same question 500 times

Tracing every entry to ComputeIndefiniteIntegral on sin(x)/(x^2 + 1)^2, which has no elementary
antiderivative and so runs the search to exhaustion:

top-level integrate calls        562
distinct top-level integrands      3

    500  sin(x) * 1 / (1 + x ^ 2) ^ 2
     50  sin(x) / (x ^ 2 + 1) ^ 2
     12  1 / (1 + x ^ 2) ^ 2 * sin(x)

Those three are one integral written three ways. And inside a single call there is a second layer
of the same thing — 5330 entries for 23 distinct integrands on e^x/(x^2 + 1)^2 — because the
solvers overlap: substitution, partial fractions, splitting a sum and by parts each decompose the
integrand differently and produce pieces that coincide.

So the answers are kept. The whole difficulty is knowing when to stop trusting them.

A settings state, not a change count

An answer depends on the ambient settings — Codomain decides whether the logarithms carry an
abs, MaxExpansionTermCount bounds the by-parts recursion — and a setting is scoped to a flow
rather than a thread, so "nothing has changed" is not something a per-thread cache may assume.

I built the counter version first, and it gave back almost none of the improvement. A scope
opened and closed leaves every setting exactly as it found it while moving a counter twice, and the
library opens such scopes constantly: one Simplify of one unevaluated integral opened 3520 of
them, every one from numeric downcasting in Number/Operators.cs. A counter therefore reports
"everything has changed" continuously and is useless here. That is measured, not predicted — it is
why SettingsState compares state.

The state of a setting is the reference to the frame on top of its stack, which is exact rather
than a hash: releasing a scope restores the very frame object that was there before it, since the
chain below a pushed frame is never rebuilt. So a comparison is a handful of reference comparisons
with no possibility of collision, a balanced open-and-close compares equal, and no enumeration of
which settings matter
is needed — which would have been the fragile way to do this.

Reads are lock-free: registering happens while the settings are being constructed, and the reader is
on the integrator's hot path.

The cap

Emptied at 4096 entries. Nothing in it expires on its own — the settings holding still is the entire
condition for keeping an answer — so without a cap a long-lived process integrating a stream of
different expressions would hold every one of them for ever. The worst measured need is 23 distinct
integrands within a call.

Verification

Full suite 9550 passed, 0 failed, 14 skipped. 9 new tests.

The three that matter are the codomain ones, because a cache like this does not fail by
returning nonsense — it fails by returning a correct answer worked out under settings the caller is
no longer standing in. 1/x integrates to ln(abs(x)) on the reals and ln(x) on the complex
plane, so it separates them in one expression; it is asked in both orders, and with a scope entered
while an answer for the same integrand is already held.

Those three were verified to bite: with the state comparison disabled they fail, and with it they
pass. A test that passes either way would have proved nothing.

There is also a guard that sin(x)/(x^2 + 1)^2 finishes under 30 s — two orders of magnitude above
what it now takes and a quarter of what it used to, so it catches the regression without being a
benchmark that flakes.

What this replaced

Three earlier explanations of the same timings, each abandoned when measured:

Closes #1156.

`Simplify` of an integral it cannot evaluate took a minute and a half. Tracing
every entry to the integrator on `sin(x)/(x^2 + 1)^2` says why: 562 top-level
calls for 3 distinct integrands, one of them asked 500 times, and those 3 are
one integral written three ways. Inside a single call there is a second layer of
it -- 5330 entries for 23 distinct integrands on `e^x/(x^2 + 1)^2` -- because
the solvers overlap, substitution and partial fractions and splitting a sum and
by parts each decomposing the integrand differently into pieces that coincide.

So the answers are kept, and the whole difficulty is knowing when to stop
trusting them. An answer depends on the ambient settings -- Codomain decides
whether the logarithms carry an abs, MaxExpansionTermCount bounds the by-parts
recursion -- and a setting is scoped to a flow rather than a thread, so "nothing
has changed" is not something a per-thread cache may assume.

SettingsState answers it by comparing what every setting reads as. A state and
not a change count, which is the part that had to be measured rather than
reasoned about: a scope opened and closed leaves every setting exactly as it
found it while moving a counter twice, and the library opens such scopes
constantly -- one Simplify of one unevaluated integral opened 3520 of them, all
four kinds from numeric downcasting in Number/Operators.cs. A counter therefore
reports "everything has changed" continuously, and the first version of this,
built on one, gave back almost none of the improvement.

The state of a setting is the reference to the frame on top of its stack, which
is exact rather than a hash: releasing a scope restores the very frame object
that was there before it, since the chain below a pushed frame is never rebuilt.
Comparing snapshots is a handful of reference comparisons with no chance of
collision, and a balanced open-and-close compares equal, which is the point.
Reads are lock-free because registering happens while the settings are
constructed and the reader is on the integrator's hot path.

The dictionary is emptied at 4096 entries. Nothing in it expires on its own --
the settings holding still is the whole condition for keeping an answer -- so
without a cap a long-lived process integrating a stream of different
expressions would hold every one of them for ever.

Measured, both arms on one machine. Simplify of the three integrands with no
elementary antiderivative: 114,851 ms to 1163, 85,683 to 1517, 51,072 to 188,
and 1/(x^4 + 1)^2 from 3811 to 9. A single cold Integrate in a fresh process is
faster too, on all nine tried: sqrt(tan(x)) 154.5 ms to 79.1, x^2*sin(x) 21.0 to
10.2, 1/(x^3 + 1) 49.8 to 31.2.

The three tests that matter are the ones on the codomain, since a cache like
this does not fail by returning nonsense but by returning a correct answer
worked out under settings the caller is no longer standing in. `1/x` integrates
to ln(abs(x)) on the reals and ln(x) on the complex plane, asked in both orders
and with a scope entered while an answer is already held. All three fail with
the state comparison disabled, which is what makes them worth having.

Full suite 9550 passed, 0 failed.

Closes #1156.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012sonx8iAspMiwRwokT1Ura
@Rafael-SOWNet
Rafael-SOWNet merged commit 0a8dff6 into master Sep 4, 2026
31 checks passed
@Rafael-SOWNet
Rafael-SOWNet deleted the integral-answers-kept branch September 4, 2026 09:49
Rafael-SOWNet added a commit that referenced this pull request Sep 9, 2026
… instead of killing the process

Integrate could recurse without bound. A stack overflow is not an exception a
caller can handle: the process aborts, and everything it had not finished is
lost. Found by work/intbench against Rubi's independent test suites, where the
run died with SIGABRT partway through and the harness refused to publish a
report from it. The trace was thousands of frames alternating
ComputeIndefiniteIntegral with SolveBySubstitution.

It is a regression. The same corpus with the same flags on the same machine:
v2.4.0 ran past the point of failure and on into the next file; the 2.5.0
release commit aborted 130 lines into it.

Why the memo added in #1157 does not already stop this, which is the part worth
writing down. SolveBySubstitution names its new variable with
Variable.CreateUnique, so every level integrates with respect to a *fresh*
variable. The key (expr, x, integrateByParts) therefore differs at every level
even when the level is the same problem renamed, and the cache never sees the
same question twice. A set of already-visited shapes would fail for exactly the
same reason: the shapes are alpha-equivalent rather than equal. A depth bound
does not care what the levels are called, which is why it is the fix.

One detail that is easy to get wrong. A null produced by running out of descent
must not be written into the memo as though it were a fact about the integrand,
because the same key reached from less deep may well be answerable. A non-null
answer is kept whatever happened elsewhere, since an antiderivative that was
found is correct however deep the search that found it went.

Declining is a legitimate answer here and a wrong one is not: an unevaluated
integral(...) says "I could not settle this", which is true, where an aborted
process says nothing at all.

#1232

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012sonx8iAspMiwRwokT1Ura
Rafael-SOWNet added a commit that referenced this pull request Sep 9, 2026
… instead of killing the process (#1233)

Integrate could recurse without bound. A stack overflow is not an exception a
caller can handle: the process aborts, and everything it had not finished is
lost. Found by work/intbench against Rubi's independent test suites, where the
run died with SIGABRT partway through and the harness refused to publish a
report from it. The trace was thousands of frames alternating
ComputeIndefiniteIntegral with SolveBySubstitution.

It is a regression. The same corpus with the same flags on the same machine:
v2.4.0 ran past the point of failure and on into the next file; the 2.5.0
release commit aborted 130 lines into it.

Why the memo added in #1157 does not already stop this, which is the part worth
writing down. SolveBySubstitution names its new variable with
Variable.CreateUnique, so every level integrates with respect to a *fresh*
variable. The key (expr, x, integrateByParts) therefore differs at every level
even when the level is the same problem renamed, and the cache never sees the
same question twice. A set of already-visited shapes would fail for exactly the
same reason: the shapes are alpha-equivalent rather than equal. A depth bound
does not care what the levels are called, which is why it is the fix.

One detail that is easy to get wrong. A null produced by running out of descent
must not be written into the memo as though it were a fact about the integrand,
because the same key reached from less deep may well be answerable. A non-null
answer is kept whatever happened elsewhere, since an antiderivative that was
found is correct however deep the search that found it went.

Declining is a legitimate answer here and a wrong one is not: an unevaluated
integral(...) says "I could not settle this", which is true, where an aborted
process says nothing at all.

#1232


Claude-Session: https://claude.ai/code/session_012sonx8iAspMiwRwokT1Ura

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Rafael-SOWNet added a commit that referenced this pull request Sep 9, 2026
…nce (#1234)

#1233 bounded the integrator's descent, which stopped it killing the process.
That was necessary and it was not sufficient: the descent branches, so 32 levels
of a cycle is still an enormous search. Measured on Rubi's independent test
suites, problems 1400-1450 took 985 seconds under the bound alone where v2.4.0
took 44. One integrand held the run for over ten minutes.

Declining at the first repeat instead: 37 seconds over the same fifty problems,
which is v2.4.0's figure and slightly under it.

The set is keyed on the integrand with its variable renamed to one canonical
name, and that renaming is the whole point. SolveBySubstitution names each new
variable with Variable.CreateUnique, so a level and the level it came from are
alpha-equivalent rather than equal. A set keyed on the integrand as written
would never see the same entry twice -- which is exactly why the memo added in
#1157 could not stop this either, and why the first fix had to be a depth bound
rather than a visited set.

DeepestDescent stays as a backstop for anything this does not model, rather than
as the mechanism.

#1232


Claude-Session: https://claude.ai/code/session_012sonx8iAspMiwRwokT1Ura

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Simplify spends up to 90s on an integral it cannot evaluate, asking the integrator the same question 500 times

1 participant