Skip to content

feat(host-core): catch up scheduled occurrences missed while the app was not running - #1212

Open
hawchou1995 wants to merge 1 commit into
vastsa:mainfrom
hawchou1995:feat/scheduled-catch-up-missed-occurrences
Open

hawchou1995 wants to merge 1 commit into
vastsa:mainfrom
hawchou1995:feat/scheduled-catch-up-missed-occurrences

Conversation

@hawchou1995

Copy link
Copy Markdown

Closes #1178

Responding to the maintainer's invitation in #1178 ("可否 pr 一下?"). Left as feat
deliberately: R6.1 exempts work maintainers direct and states that relabelling another type
as fix/perf does not qualify it (SPEC 06-delivery/03-ai-development-workflow.md R6.1).

What

Automatic scheduling now catches up an occurrence missed while the app was not running.

Why

due() rearmed anything more than 90 seconds late into the future, and recover() rearmed
every armed task at startup. The intent was to avoid a burst of stale runs after downtime, but
the effect is that closing the app — or letting the machine sleep — across a scheduled minute
loses that run silently and permanently. Both sites are changed to admit one catch-up when the
task's newest miss is still inside its window, and to rearm as before when it is not.

Boundaries — the answer to "最大重试天数或次数?"

A retry count answers the wrong question. For a recurring task only the newest miss is worth
running, because later occurrences supersede the older ones: replaying n daily occurrences is
not n retries of one run, it is n stale runs. So the policy is a window, not a count:

Key Default Notes
catchUp true per task
catchUpWindowMinutes 180 hourly, 1440 daily/weekly clamped to 5–10080
  • One miss produces at most one run; the dispatch path's existing rearm keeps it from repeating.
  • Only the newest miss is considered, so 30 days offline still yields a single catch-up.
  • Anything older than the window is rearmed into the future instead of replayed.
  • At most two catch-ups are admitted per poll, newest first; the overflow waits for a later poll.
  • Unchanged invariants: paused, unarmed, manual cadence and a task with an unfinished run never
    catch up, and a catch-up goes through the same scheduled.run path as an automatic occurrence,
    so the run ledger, overlap suppression and SCHEDULE_NOT_DUE admission all still apply.

Both keys ride the task's existing config_json boundary, so there is no new table, column,
migration, RPC method or wire field.

Scope

  • crates/host-core/src/scheduled/automation.rs — catch-up policy, rewritten due()/recover(), tests
  • crates/host-core/src/scheduled/timing.rs — Schedule::latest_missed() and its tests
  • Docs: new ADR 0310, the ADR index, the amended automation ADR, guide/automations, spec/04-ux/01-ui-ia
    and the decisions log (D635), English and Simplified Chinese in step
  • No dependency, Cargo.toml, Cargo.lock, schema, RPC or JS/TS change. reschedule, running
    and begin_run keep their signatures and semantics.
  • No Scheduled-page form controls in this change; the defaults apply without configuration and the
    existing Agent tools can set both keys per task. A follow-up can add controls.

Verification

Local Windows candidate, rustc/cargo 1.98.1 (x86_64-pc-windows-gnu), branch head 6dedf39f on
origin/main 7f6c2cd9:

Check Result
cargo test -p host-core --locked 670 passed, 3 failed
Same, excluding the three pre-existing environment failures 670 passed, 0 failed
cargo test -p host-core --locked scheduled 45 passed, 0 failed (12 of them new)
cargo fmt --all -- --check clean
cargo clippy -p host-core --all-targets clean, no warnings
node docs/scripts/check-docs.mjs 531 pages verified
node docs/scripts/check-locales.mjs 83 English/Chinese pairs verified
node scripts/check-architecture.mjs passed (Rust 511 / 439 LOC, limit 1000)
node scripts/check-pr-base-main.mjs passed: origin/main is an ancestor of HEAD

The three failures are not caused by this change: the unchanged baseline fails the same three
with the same assertion, CONFIG_SYNC_LIMIT_EXCEEDED: instruction file is too large. Root cause is
that config_sync::domains::global_instruction_path() reads dirs::home_dir()/.pi/agent/AGENTS.md
from the real home directory rather than an isolated temp root, and that file is 33,559 B here
against MAX_INSTRUCTION_BYTES = 32 KiB — the repository's own AGENTS.md (33,277 B) is over the
same limit. A clean CI home directory has no such file.

New tests cover: catch-up once inside the window; rearm outside it; catchUp: false; manual and
unarmed tasks; paused tasks; a running task suppressing its own catch-up; the two-per-poll cap;
startup keeping an in-window miss for exactly one catch-up; the default windows and the clamp;
legacy tasks never catching up; and the calendar arithmetic (daily/weekly collapse to the newest
occurrence, hourly anchored at the armed instant, invalid input) against an injected timezone.

NOT RUN

  • pnpm test:e2e:scheduled (E2E-SCHEDULED-desktop-automation-lifecycle). That suite runs the real
    desktop, preload, Host SQLite and agent sidecar, so it needs a built Electron candidate; this
    candidate has no node_modules or desktop build, and installing/building to produce one would
    write several GB to a system drive that is already near full. Reason recorded rather than guessed.
    Alternative validation: the scenario's Expected clause delegates stale/missed occurrence handling
    to Host tests — "Host tests additionally prove duplicate admission rejection, stale/missed
    occurrence handling, invalid input rejection and recovery" — which is the path extended here and
    is covered by the 12 new Host tests above. Remaining risk: the renderer-facing behaviour of a
    catch-up run (run-history row, conversation link, no foreground focus change) is unverified end
    to end, though it reuses the existing automatic dispatch path unchanged.
  • pnpm install, pnpm build:js, pnpm lint, pnpm -r test. No JS/TS file is touched, and the
    two dependency-free docs checkers were run instead (above). CI covers the rest.

Compatibility

Merging is additive: existing tasks keep their cadence, nextRunAt semantics and run history, and
gain catch-up with defaults. Reverting restores the previous behaviour, and a task carrying
catchUp / catchUpWindowMinutes while reverted is unaffected — they are two keys the older code
ignores in config_json, so no data cleanup is needed.

Affected specs and records: ADR 0310 (new), ADR scheduled-desktop-automations (amended),
ADR 0305 (unchanged), 04-ux/01-ui-ia §3.3, guide/automations, decisions log D635,
E2E-SCHEDULED-desktop-automation-lifecycle.

…was not running

Automatic dispatch only fires while PI-Desktop is running, and the scheduler
deliberately dropped anything it found late: due() rearmed an occurrence more
than 90 seconds old into the future, and recover() rearmed every task at
startup. Closing the app across a scheduled minute therefore lost that run
silently and permanently.

Admit the single newest miss instead, while it is still inside the task's
catch-up window. The window is the boundary rather than a retry count, because a
recurring task's later occurrences supersede the older ones: one miss produces at
most one run, at most two catch-ups start per poll, and anything older is rearmed
as before. Paused tasks, tasks without a schedule, manual cadence and a task
with an unfinished run never catch up, and a catch-up is dispatched through the
same scheduled.run path as an automatic occurrence.

catchUp (on by default) and catchUpWindowMinutes (three hours hourly, one day
daily and weekly, clamped to 5-10080) ride the task's existing config_json
boundary, so there is no schema, RPC, wire or dependency change.

Refs vastsa#1178

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Feature] 定时任务支持上线后延时触发

1 participant