Skip to content

fix(runtime): retry mid-turn compaction after a no_safe_completed_span miss - #5792

Open
Colafornia wants to merge 4 commits into
apache:mainfrom
Colafornia:fix/5790-compaction-no-safe-span
Open

Colafornia wants to merge 4 commits into
apache:mainfrom
Colafornia:fix/5790-compaction-no-safe-span

Conversation

@Colafornia

Copy link
Copy Markdown
Member

Summary

Fixes #5790.

A mid-turn compaction attempt that finds nothing safe to fold (no_safe_completed_span) never reaches the summarizer, yet it latched state.summarizerFailure for the rest of the turn. Every later attempt short-circuited before re-reading the ledger — including the reactive recovery meant to rescue a real provider context_overflow — so a turn could die on input capacity even though a grown pool would have folded.

The latch now covers only outcomes that actually invoked the summarizer (provider_error, malformed_*, output_length, input_too_large, empty summary), preserving the bounded-retry behavior added for #4634. no_safe_completed_span is a property of the event pool at that step, so the next attempt re-reads the ledger and can fold once a completed tool pair appears.

Verification

  • node --test on the rebuilt @maka/runtime dist: mid-turn-capacity-backend.test.js, overflow-reactive-recovery.test.js, history-compaction.test.js — 157 tests pass.
  • Two new tests reproduce the issue deterministically through the backend fixtures (the proactive trigger and the reactive recovery path); reverting the fix makes exactly those fail.
  • npm run lint and npm run format:check pass repo-wide; tsc passes via npm --workspace @maka/runtime run build.
  • Not run: the full npm test matrix across all workspaces, knip, and the desktop e2e budget check.

AI use

  • Generative tooling made a substantive contribution

Tool(s) and scope: Devin authored the fix and the tests.

Checklist

  • Tests cover the change and fail without it
  • Lint, format, typecheck and the affected suites pass locally

Does this PR entail a change in behavior?

  • Yes — described under Summary above
  • No
中文版本

摘要

修复 #5790。

一次轮中压缩尝试如果找不到可安全折叠的区间(no_safe_completed_span),并没有发起任何 summarizer 调用,却把 state.summarizerFailure 置位到本轮结束。之后的每次尝试都在重读账本前被短路——包括本应用于抢救真实 provider context_overflow 的响应式恢复——因此即使增长后的事件池本来可以折叠,turn 仍可能死于输入超限。

现在只有真正调用过 summarizer 的结果才置位(provider_error、malformed_*、output_length、input_too_large、空摘要),#4634 引入的有界重试行为不变。no_safe_completed_span 只是该步事件池的属性,下一次尝试会重读账本,在出现完整的工具调用对后即可折叠。

验证

  • 在重新构建的 @maka/runtime dist 上运行 node --test:mid-turn-capacity-backend.test.js、overflow-reactive-recovery.test.js、history-compaction.test.js —— 157 个测试通过。
  • 新增两个测试,通过后端 fixture 确定性复现该问题(主动触发路径与响应式恢复路径);还原修复代码后恰好这两个测试失败。
  • npm run lint 与 npm run format:check 仓库级通过;npm --workspace @maka/runtime run build 的 tsc 通过。
  • 未运行:全 workspace 的 npm test 矩阵、knip、desktop e2e 预算检查。

…n miss

state.summarizerFailure latched on every fail-open reason, including
no_safe_completed_span — an outcome that never reaches the summarizer and
only describes the event pool at that step. Once set, it short-circuited
every later attempt for the rest of the turn, including the reactive
overflow recovery that could have folded the grown ledger, so the turn
could die on input capacity.

Latch only outcomes that actually invoked the summarizer
(provider_error, malformed_*, output_length, input_too_large, empty
summary): the bounded-retry behavior from apache#4634 is unchanged, while a
pool that grows a completed tool pair is re-evaluated on the next
attempt.

Fixes apache#5790

Generated-by: Devin

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@github-actions github-actions Bot added the effort/M Under 500 readable lines label Sep 28, 2026
@Colafornia
Colafornia marked this pull request as ready for review September 28, 2026 14:17

@jackwener jackwener left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed head 3ca4ea16cab22445a51b4108ab3c3cd9a54de7a2. The diagnosis is right: a plain no_safe_completed_span never reaches the summarizer, so it should not spend the Turn-wide latch. plan.reason has only two fail-open values (summarizer_failed, no_safe_completed_span), so the new condition is exhaustive. 1×P2 (inline).

P2: after the one-step retreat, no_safe_completed_span can follow a real summarizer call, and the latch no longer catches it.

  • In planHistoryCompaction (history-compaction.ts, around 380-395), a summarizer rejection with a proven accepted-input boundary sets maxCoveredCount = proven and continues.
  • If the smaller window then has no safe coverage, the same call returns { reason: 'no_safe_completed_span' } (:316, or :464 when the loop runs out) even though the summarizer was already dispatched and failed.
  • With this change that outcome does not latch. So every later step re-plans the same prefix, repeats the same doomed summarizer call, and retreats into the same miss. That is the repeated-summarizer-call loop #4634 bounded, and a slow provider that fails makes each iteration expensive.
  • Suggested fix: have the planner report that it invoked the summarizer, and latch on that rather than on reason. Alternatively, return summarizer_failed (with the original diagnosticReason) when the retreat itself finds no safe span. Please add a test: an input_too_large rejection whose proven boundary leaves no safe coverage, then a second step, asserting that no second summarizer call is made.

The new tests cover the proactive trigger and the reactive-recovery path for the plain miss, and CI test is green. Not run locally: the suites. The P2 comes from reading the code paths above; I did not reproduce it end to end.

Comment thread packages/runtime/src/ai-sdk-compaction.ts
Colafornia and others added 2 commits September 28, 2026 22:57
When an input_too_large rejection retreats to the proven boundary and the
smaller window has no safe coverage, the summarizer was already dispatched
and failed, so the result must not look like a pool that never reached it.
Return summarizer_failed with input_too_large for that miss so the turn
latch catches it instead of re-dispatching the same doomed call every step.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…overage

The proven boundary a retreat targets is a prior-run reply, which always
precedes the head anchor that mid_turn coverage must cover past — so the
retreated cut could never fold and the miss disguised the summarizer
failure. Mid_turn now reports input_too_large directly; the retreat and its
post-retreat miss handling remain only where they can succeed
(standalone/pre_turn).

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

@hqhq1025 hqhq1025 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed commit 826b419. The change lets a structural no_safe_completed_span retry when the event pool grows, while retaining the Turn-wide latch after a real summarizer failure. A later commit removes all mid-turn input_too_large retreats. I found one P2 regression in that removal, inline.

Node 24 build:test and the three focused compaction suites passed (160 tests). I also exercised the planner with a prior-run, same-turn model reply after the head anchor: it made one summarizer call and returned summarizer_failed without testing the proven smaller prefix. Fresh main 71bc045 merges cleanly and git diff --check passes. Current-head hosted test was queued at review time. I did not run the full suite or a real provider/handoff process; the concrete handoff path and synthetic rejection establish the reachable boundary, not its production frequency. Please address the inline regression before merge. This is not merge approval.

Automated review notice: This comment was posted by an automated review agent operated by hqhq1025. It is not an independent human review and does not replace one.

Comment thread packages/runtime/src/history-compaction.ts Outdated

@jackwener jackwener left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving at 826b41976134c2a2a80840d411a78993b88df1f0. The P2 from review 5339998409 is closed, and no P0–P2 remain.

  • a6b6c2227: once a retreat has happened (retreated = true), a smaller window with no safe coverage, or a loop that runs out, now returns summarizer_failed / input_too_large instead of no_safe_completed_span. The Turn latch therefore catches a failure the summarizer actually produced. A plain miss that never reached the summarizer still returns no_safe_completed_span, which is #5790's fix.
  • 826b41976: mid_turn no longer retreats. This checks out: the mid-turn planner receives state.priorInvocations, so acceptedInputBoundary can only land on a prior-run reply, which precedes the head anchor that mid_turn coverage must pass. The retreated cut could never fold. input_too_large now fails open as summarizer_failed directly and latches. standalone and pre_turn keep the retreat.
  • New tests cover the retreat-then-miss case in the planner and the mid-turn backend path.

CI test is green on this head.

Not run locally: the suites.

Handoff replay keeps the same logical turn's predecessor events and run
record in the pool, so an on-route reply after the head anchor can prove a
retreat boundary. Skipping the retreat for all mid_turn folds dropped that
reachable fold and spent the Turn latch; restore the retreat and cover the
handoff-shaped boundary in the planner tests.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

@hqhq1025 hqhq1025 left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed commit fa32a3f. The new delta removes the unconditional mid-turn failure on input_too_large (packages/runtime/src/history-compaction.ts:389-405) and restores the one-step proven-boundary retreat. The added planner regression covers a handoff predecessor reply after the current Turn's head anchor, asserting a second, smaller coverage attempt; the retreat-then-no-safe-span case still reports summarizer_failed, so the Turn latch bounds repeated calls. The earlier P2 is addressed. No new substantiated P0-P3 in the inspected change. Node 24 build:test and three focused compaction files passed (161 tests); current-head hosted test is green. Fresh main 2f32205 merges cleanly and git diff --check passes. I did not run the entire local suite or a real provider/handoff process. This is not merge approval.

Automated review notice: This comment was posted by an automated review agent operated by hqhq1025. It is not an independent human review and does not replace one.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

effort/M Under 500 readable lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

bug(runtime): "nothing to fold" disables compaction for the turn

3 participants