Skip to content

The integrator's descent is bounded, so a substitution cycle declines instead of killing the process - #1233

Merged
Rafael-SOWNet merged 1 commit into
masterfrom
integrator-descent-bound
Sep 9, 2026
Merged

Rafael-SOWNet merged 1 commit into
masterfrom
integrator-descent-bound

Conversation

@Rafael-SOWNet

Copy link
Copy Markdown
Member

Fixes #1232.

Integrate could recurse without bound and take the process down. A stack overflow is not an exception a caller can handle: the process aborts and everything it had not finished is lost.

Found while running the release checklist for 2.5.0 — work/intbench against Rubi's independent test suites died with SIGABRT partway through, and the harness correctly refused to publish a report from a run that had not finished.

It is a regression, and that was measured rather than assumed

Same corpus, same flags (--families=0 --per-file=100000 --budget=5), same machine:

build outcome
v2.4.0 (de7189b1) completed — 599/1774 solved, 0 wrong, 58 timeouts
2.5.0 release candidate aborted partway through

Why the memo added in #1157 does not already stop it

This is the part worth writing down, because it is the reason the obvious fix is the wrong one.

SolveBySubstitution names its new variable with Variable.CreateUnique:

var uSub = Variable.CreateUnique(expr, "u_sub");
...
Integration.ComputeIndefiniteIntegral(integrandInU, uSub, integrateByParts)

Every level therefore integrates with respect to a fresh variable. The memo key (expr, x, integrateByParts) differs at every level even when the level is the same problem renamed, so the cache never sees the same question twice and cannot break a cycle. A set of already-visited shapes fails for exactly the same reason: the shapes are alpha-equivalent rather than equal.

A depth bound does not care what the levels are called, which is why that is what this does.

The fix

A bound on the descent, and a refusal when it is reached. Declining is a legitimate answer here and a wrong one is not: an unevaluated integral(...) says "I could not settle this", which is true, where an aborted process says nothing at all.

The number is measured, and my first guess at it was wrong in an instructive way. I started at 64 on the reasoning that a bound only has to stop the overflow. It did stop it — and left the corpus run crawling, because this descent branches: the existing remark on the memo records a single call entering the integrator 5,330 times for 23 distinct integrands. Depth that is never legitimately used still gets searched before it is abandoned, so a loose bound is barely a bound.

So I instrumented it and measured how deep answered integrals actually go, over twenty-three integrands from the hard end of the corpus:

depth integrand
13 x^5*cosh(x)
8 x^2*sqrt(5-x^2)
7 e^(x^(1/3))
6 x^3/(1+x)^10, x^5*sqrt(1+x^2)
≤ 4 the other eighteen

32 is about two and a half times the deepest real descent seen, and the measurements are in the code so the next person does not have to re-derive them.

One detail that is easy to get wrong, and is called out in the code: a null produced by running out of descent must not be written into the memo as though it were a fact about the integrand, because the same key reached from less deep may well be answerable. A non-null answer is kept whatever happened elsewhere, since an antiderivative that was found is correct however deep the search that found it went.

What I could not do, stated rather than glossed

I could not isolate a single integrand that reproduces this standalone. Several candidates from the neighbourhood of the failure — ln(tan(x))/(cos(x)*sin(x)), x^5*cosh(x), e^(x^(1/3)), 1/(4+x+sqrt(1+x)), x^3/(1+x)^10 — all return promptly when integrated on their own, on the unfixed build.

The reason is in the harness's own design note: a case runs on its own thread so that a hang costs the budget rather than the run, and .NET cannot abort a thread. So a runaway is abandoned at five seconds while its thread keeps recursing in the background, and the overflow that kills the process happens later, on a thread nobody is waiting for, unrelated to whichever problem is being printed at the time. That is also why re-running it does not abort at the same problem index.

So there is no unit test here. The guard is the harness, and the evidence is that the harness now completes.

🤖 Generated with Claude Code

https://claude.ai/code/session_012sonx8iAspMiwRwokT1Ura

… instead of killing the process

Integrate could recurse without bound. A stack overflow is not an exception a
caller can handle: the process aborts, and everything it had not finished is
lost. Found by work/intbench against Rubi's independent test suites, where the
run died with SIGABRT partway through and the harness refused to publish a
report from it. The trace was thousands of frames alternating
ComputeIndefiniteIntegral with SolveBySubstitution.

It is a regression. The same corpus with the same flags on the same machine:
v2.4.0 ran past the point of failure and on into the next file; the 2.5.0
release commit aborted 130 lines into it.

Why the memo added in #1157 does not already stop this, which is the part worth
writing down. SolveBySubstitution names its new variable with
Variable.CreateUnique, so every level integrates with respect to a *fresh*
variable. The key (expr, x, integrateByParts) therefore differs at every level
even when the level is the same problem renamed, and the cache never sees the
same question twice. A set of already-visited shapes would fail for exactly the
same reason: the shapes are alpha-equivalent rather than equal. A depth bound
does not care what the levels are called, which is why it is the fix.

One detail that is easy to get wrong. A null produced by running out of descent
must not be written into the memo as though it were a fact about the integrand,
because the same key reached from less deep may well be answerable. A non-null
answer is kept whatever happened elsewhere, since an antiderivative that was
found is correct however deep the search that found it went.

Declining is a legitimate answer here and a wrong one is not: an unevaluated
integral(...) says "I could not settle this", which is true, where an aborted
process says nothing at all.

#1232

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012sonx8iAspMiwRwokT1Ura
@Rafael-SOWNet
Rafael-SOWNet merged commit 229c9ab into master Sep 9, 2026
31 checks passed
@Rafael-SOWNet
Rafael-SOWNet deleted the integrator-descent-bound branch September 9, 2026 14:33
Rafael-SOWNet added a commit that referenced this pull request Sep 9, 2026
…nce (#1234)

#1233 bounded the integrator's descent, which stopped it killing the process.
That was necessary and it was not sufficient: the descent branches, so 32 levels
of a cycle is still an enormous search. Measured on Rubi's independent test
suites, problems 1400-1450 took 985 seconds under the bound alone where v2.4.0
took 44. One integrand held the run for over ten minutes.

Declining at the first repeat instead: 37 seconds over the same fifty problems,
which is v2.4.0's figure and slightly under it.

The set is keyed on the integrand with its variable renamed to one canonical
name, and that renaming is the whole point. SolveBySubstitution names each new
variable with Variable.CreateUnique, so a level and the level it came from are
alpha-equivalent rather than equal. A set keyed on the integrand as written
would never see the same entry twice -- which is exactly why the memo added in
#1157 could not stop this either, and why the first fix had to be a depth bound
rather than a visited set.

DeepestDescent stays as a backstop for anything this does not model, rather than
as the mechanism.

#1232


Claude-Session: https://claude.ai/code/session_012sonx8iAspMiwRwokT1Ura

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Integrate recurses without bound and takes the process down, and it is a 2.5.0 regression

1 participant