Skip to content

[Bug] 摘要失败后立刻丢弃旧上下文,不会重试压缩直到成功 / Failed compaction drops history instead of retrying summarization #543

Description

@YahooYuan666

What happened? / 发生了什么

Automatic context compaction often does not produce a real summary. On failure it immediately writes a retained_tail fallback checkpoint (~112 tokens of recovery prose) and omits older messages from the next model request. The UI still shows this as a successful row: 上下文已压缩 · 第 1 次 摘要 ≈ 112 tokens.

This is distinct from #224. #224 / PR #249 made the fallback recoverable (non-empty last-user tail, strip the fallback marker so a later compaction can try a real summary). On this compaction attempt, summarization is still one-shot: fail → drop history. There is no retry-until-a-real-summary-exists loop.

长会话自动压缩经常压不出真正摘要。失败后立刻写入 retained_tail 回退检查点(约 112 tokens 的恢复说明),并从下一次模型请求里丢掉旧消息。UI 仍显示成成功:上下文已压缩 · 第 1 次 摘要 ≈ 112 tokens。

这和 #224 不是同一件事。#224 / PR #249 让回退可恢复(非空的最后一条用户消息、剥掉 fallback 标记以便下一次压缩再试真摘要)。这一次压缩仍然是一锤子:失败 → 丢历史。没有「重试直到压出真摘要」的环。

Steps to reproduce / 复现步骤

  1. Open PI-Desktop 0.14.8 on Windows, Agent mode, a long coding session.

  2. Let the context grow until automatic compaction fires (hard limit). Typical tokensBefore here: ~190k–330k, one case ~885k.

  3. Watch the transcript row 上下文已压缩 · 第 N 次 摘要 ≈ 112 tokens.

  4. Ask the model about decisions, file paths, or unfinished work from before that row.

  5. Windows 上打开 PI-Desktop 0.14.8,Agent 模式,跑长会话。

  6. 等到自动压缩触发(硬上限)。本机 tokensBefore 常见约 19–33 万,有一次约 88 万。

  7. 看到 上下文已压缩 · 第 N 次 摘要 ≈ 112 tokens。

  8. 再问压缩点之前的决策、路径、未完成工作。

Expected behavior / 预期行为

Compaction is a model summarization pass over the conversation.

  • Transient failure (timeout, 5xx, dropped stream): retry the same summary request until it completes.
  • Input too large for one request: do not skip the model. Shrink or chunk (map-reduce / truncate tool output) and retry until a real structured summary exists.
  • Only after retries are exhausted should a safety net keep a bounded recent original tail (on the order of keep-recent, not one user message).
  • The UI must not present a failure stub as 摘要 ≈ 112 tokens. Say that summary generation failed.

压缩的本质就是调模型把上下文走一遍总结。

  • 瞬时失败(超时、5xx、断流):重试同一次摘要请求直到完成。
  • 一次塞不下:不要跳过模型。缩小或分段(map-reduce / 截断工具输出)再压,直到产出结构化真摘要。
  • 只有重试耗尽,才允许保有界近期原文(keep-recent 量级,而不是一条用户消息)。
  • UI 不要把失败占位显示成 摘要 ≈ 112 tokens,应标明摘要生成失败。

Actual behavior / 实际行为

Inspected local session JSONL + installed resources/agent-runtime/sidecar.js on 0.14.8:

  1. Codex-shaped preparation folds all history into messagesToSummarize and keeps at most the latest user message (active_turn) or none (completed_turn). keepRecentTokens no longer decides what survives.
  2. compactionSummaryWouldExceedBudget — if history + previous summary ≥ window − output budget − 2048 — never calls the model and fails immediately.
  3. generateCompaction / compact() may retry transient stream drops internally, but buildCheckpoint does not re-run summarization with a smaller payload.
  4. recoverCompactionFailure writes failureCode: CONTEXT_COMPACTION_FAILED, fallback: retained_tail, summary ≈:
No previous context checkpoint is available.

[automatic context recovery: older context was omitted after summary generation failed]
…
Older messages before this checkpoint are omitted from the next model request.

That stub is ~451 chars ≈ 112 tokens — exact match for the UI.

  1. retainedTailForContext then keeps only the latest user message, or [] on completed_turn.

Local counts on this machine (sanitized): 6 / 10 compaction entries were this stub. The 4 successes produced real 4–6k-character summaries (~1k–1.5k tokens). So when summarization runs, it works; the failure path is what wipes memory.

本机 0.14.8 的 session JSONL 和安装包 sidecar.js:

  1. Codex 形状把全部历史折进 messagesToSummarize,最多留最后一条用户消息(active_turn),结束回合则一条不留。keepRecentTokens 不再决定谁活下来。
  2. compactionSummaryWouldExceedBudget:历史 + 上次摘要 ≥ 窗口 − 输出预算 − 2048 时,不调模型,直接失败。
  3. compact() 内部可能重试断流,但 buildCheckpoint 不会缩小输入再跑一轮摘要。
  4. recoverCompactionFailure 写入上述 ~112 tokens 占位。
  5. retainedTailForContext 只留最后一条用户消息,或 completed_turn 时为空。

本机 10 次压缩里 6 次是这种占位;4 次成功摘要有 4–6k 字。模型能压时压得出,失败路径才会把记忆清掉。

App version / 应用版本

0.14.8

Operating system / 操作系统

Windows

Extra environment / 额外环境信息

Logs / 日志

Compaction checkpoint details (typical failure):

{
  "failureCode": "CONTEXT_COMPACTION_FAILED",
  "fallback": "retained_tail",
  "generation": 1,
  "retainedTailMode": "active_turn"
}

tokensBefore examples: 198218, 205803, 329313, 885339.

No full transcript attached (privacy). Happy to add redacted sidecar traces if useful.

Related

Suggested fix (for maintainers)

In packages/agent-runtime compaction path:

  1. On transient summary failure: retry the summary call with backoff (cap N).
  2. On compactionSummaryWouldExceedBudget: chunk or truncate tool results and still call the model; merge partial summaries.
  3. Last resort only: persist a bounded recent original tail (8k–20k), not a 112-token notice.
  4. UI: if failureCode === CONTEXT_COMPACTION_FAILED, show “摘要生成失败”, not “摘要 ≈ 112 tokens”.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions