Skip to content

Release 0.1.25: verification and consumer adoption #192

Description

@DorianZheng

TL;DR

Prepare 0.1.25 by correcting misleading lifecycle tests and oversized prompts, then verify installation separately from the release merge.

Scope

Resolve #176, #189, and #190 without weakening approval, publication, or timeout contracts. Validate the shared auditor reminder on a live supported host. Preserve held consumers and unrelated worktree changes during adoption.

Related work and lessons

  • The FIFO dossier case already distinguishes bounded file rejection from the following synchronous audit. Apply that distinction to request metadata; a two-second whole-audit alarm conflates phases.
  • The watcher signal case polls for startup but signals even when readiness was never observed. Require the actual startup event before asserting shutdown behavior. A controlled three-second startup delay exposes the missing precondition.
  • Existing prompt budgets bound complete documents. Native timed-question plus PR recovery rendering measures 1,249 bytes against 1,200. Shorten duplicated prose while retaining evidence, human approval, and fail-closed semantics; do not enlarge limits.
  • Bats descriptor-lifetime guidance distinguishes process completion from inherited descriptor lifetime. Keep child cleanup explicit; importing a test framework would not solve these boundary mistakes.
  • Monitor documentation bounds single-prompt monitors to ten minutes. Verify actual tool delivery and dossier completion separately; fixture assertions alone cannot establish model compliance.

How it works

Keep the two-second FIFO replacement check; supervise completion of the subsequent audit separately. Require watcher startup before sending TERM, with a bounded readiness failure and explicit cleanup. Keep terminal-event/orphan assertions intact.

Reduce commit criteria and PR acknowledgment prose inside existing budgets. Preserve source validation, uncertainty findings, advisory separation, native question routing, and exact human replies. No writing-skill changes or fallback scheduler.

Validation

Reproduce existing failures and controlled delayed-audit/startup cases first. For production prompt edits, restore all non-test files to baseline and observe the existing public-boundary failures, then restore the complete changes and rerun. Run syntax, ShellCheck, focused suites, all plugin suites, host parity, architecture, installer/bootstrap, and guidance checks. Record live-host limitations explicitly.

Delivery

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions