Skip to content

feat(handoff): dual-form context thresholds + model-preset registry - #7

Merged
DumoeDss merged 4 commits into
dev/0.1.4from
feat/handoff-model-presets
Jul 16, 2026
Merged

DumoeDss merged 4 commits into
dev/0.1.4from
feat/handoff-model-presets

Conversation

@DumoeDss

Copy link
Copy Markdown
Owner

Why

All context-occupancy thresholds (handoff default 0.5, reuse default 0.25, per-role overrides) are fractions of the context window. For 1M-context models a fraction works, but for small-window models (e.g. the GPT-5.6 family at ~300K) a fraction fires far too early in absolute-token terms — and those models degrade gracefully near their limit, so the remaining absolute headroom is the better signal. Configs need an absolute-token threshold form, and the CLI should ship sensible per-model presets so users of small-window models get good behavior without hand-tuning.

What Changes

  • Handoff and reuse threshold values (pipeline-level, per-role, stage-level) accept a second form: { remainingTokens: <positive integer> } — an absolute required-headroom threshold — alongside the existing bare fraction in (0, 1]. A bare number is ALWAYS a fraction; the absolute form is ALWAYS the object, so no value is ambiguous.
  • A built-in model-preset registry (src/core/model-presets.ts) keyed by model-id substring patterns provides, per known model family: the context-window size and suggested handoff/reuse threshold values. It subsumes the existing ad-hoc model→limit map in resolveModelLimit.
  • Threshold resolution gains a preset layer just above built-in defaults: stage handoff > pipeline handoff.roles[<role>] > pipeline handoff > model preset (via the stage's resolved model) > built-in default. Reuse: reuse.roles[<role>] > reuse.threshold > model preset (via agents[<role>] model) > built-in default. Any ordinary config value therefore overrides a preset.
  • rasen pipeline show --json reports resolved thresholds in whichever form they resolved to (bare number for fractions — byte-identical for existing configs — or the { remainingTokens } object), and source can now be preset.
  • rasen agent context output gains remainingTokens (limit − contextTokens) so absolute thresholds can be compared directly against a probe.
  • The orchestration playbook and handoff templates state the comparison rule for both forms (fraction: pct >= t hands off / pct <= t allows reuse; absolute: remainingTokens <= N hands off / remainingTokens >= N allows reuse).
  • Backward compatible: existing fraction configs parse and resolve byte-for-byte identically; no version bump.

Review

Went through 3 rounds of an independent review-cycle loop — clean, with all Major findings resolved (test-tautology fixes for preset ordering and byte-identity guarantees, and a source provenance fix so pipeline show never mislabels a preset-sourced threshold as coming from a config layer). Full findings in the change's review-report.md.

Test plan

  • Full suite: 134 files / 2796 tests, green (after merging in dev/0.1.4's concurrent fix/codex-host-compat merge and rebuilding dist/)
  • tsc --noEmit clean
  • rasen validate handoff-absolute-thresholds — valid
  • Manual smoke: a fixture pipeline with handoff: { threshold: { remainingTokens: 60000 } } validates and shows correctly; a fixture with agents.implementer.model: gpt-5.x and no thresholds resolves source: preset

🤖 Generated with Claude Code

https://claude.ai/code/session_011an7Av8HWb8tWFrUPD9ZTi

DumoeDss and others added 4 commits July 16, 2026 02:20
"Control the ideas, not the code." — rasen is presented as an engineered
outer loop around the agent's inner loops; spec is heritage/internal
mechanism, not the user's input burden. Aligns docs/README.md,
docs/faq.md, docs/reviewing-changes.md, docs/getting-started.md and
their zh counterparts with the rasen.io site copy.

Ref: rasen change site-messaging-seo-geo
Fraction thresholds fire too early in absolute-token terms for small-window
models. Handoff/reuse thresholds now accept either a bare fraction (0, 1] or
an absolute `{ remainingTokens: N }` headroom form, resolved through a new
built-in model-preset registry (src/core/model-presets.ts) slotted just above
built-in defaults in the precedence chain — any configured value still wins.
`rasen agent context` and `pipeline show --json` surface the new
`remainingTokens`/preset-sourced fields; existing fraction-only configs
resolve byte-for-byte identically.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011an7Av8HWb8tWFrUPD9ZTi
…el-presets

# Conflicts:
#	src/core/agent-context.ts
#	test/core/templates/skill-templates-parity.test.ts
@DumoeDss
DumoeDss merged commit cca3427 into dev/0.1.4 Jul 16, 2026
1 check passed
DumoeDss added a commit that referenced this pull request Jul 16, 2026
…er to dual-form thresholds

Semantic merge with upstream PR #7 (dual-form handoff thresholds + model-preset
registry): merged precedence chain is now stage > role > pipeline >
project-config > global-config > preset > default, with threshold-specific
source attribution across all seven layers. The unified config surface
(key registry, effective-config, CLI set/editor, HTTP API, web UI) now accepts
both threshold forms — bare fraction (0,1] or { remainingTokens: N } — with a
shared exported thresholdSchema() validator, dual-direction shouldHandoff,
remainingTokensGt wire constraint, and a form-toggle control in packages/ui.

Verification: root tsc clean, 2974/2974 root tests, ui 59/59, live smoke of
both forms via temp RASEN_HOME; merge-reviewed clean (0 Blocker / 0 Major,
3 Minors fixed in-loop, non-author verified).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017BLKfwBAtTHUFFrgjmfmkT
@DumoeDss
DumoeDss deleted the feat/handoff-model-presets branch July 20, 2026 09:22
DumoeDss added a commit that referenced this pull request Jul 27, 2026
… declaration is locator-only (PR#88 M6)

storePermitsProject now returns true iff resolveProjectMembership is non-null (Store project record). The Project-side Store declaration no longer grants Session eligibility on its own — it is a locator only, per locked decisions §33 (#7/#8/#9). The 5 declaration helpers stay imported and are reused ONLY to classify the rejection diagnostic:
- declaration names this Store but no record -> 'legacy declaration-only install' marker + 'rasen store add-project' repair
- declaration absent / points elsewhere -> plain missing-record message + same repair
No new wire code enum (marker is a message substring); no auto-grant; no auto-rewrite. Discovery (project->store, listProjectStoreCandidates) stays a union — only store->project eligibility became record-only.
Spec deltas: session-runtime-context + store-project-membership (titles match canonical verbatim; disjoint from C2's two titles).
Review: clean, 0 findings. 33 tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
DumoeDss added a commit that referenced this pull request Jul 30, 2026
…..5) + archive (#111)

* feat(ecp-review-cycle): import in-flight ReviewCycle vertical foundation

Imports the uncommitted ECP-1 front-half + domain model written in a prior
session, as the baseline for the feat/ecp-review-cycle branch:

- Domain reducer (src/core/change-run/internal/review-cycle.ts): Zod-validated
  review/triage/fix/re-review contracts, event state machine, actor separation,
  ship guard, round-cap exhaustion. 6 domain tests pass.
- Runtime adapter (review-cycle-runtime.ts): projection + pre-commit validation
  (orphaned pending back-half wiring; 3 expected TS errors).
- Front half: definition prepare (supportsV2ReviewCycleRuntime), support analysis
  (supported_v2_review_cycle), lowerer (lowerV2ReviewCyclePlanInput ->
  bounded-loop RuntimePlanNode), runtime-plan types + validateBoundedLoop.
- Direction docs: executable-composite-pipelines sub-direction (target-state,
  roadmap, slice spec/plan/result), research relocation, parent roadmap ECP
  recalibration snapshot, architecture doc, 0.1.6 completion audit.

Back half (reconciler bounded-loop execution, record enrichment, reducer/facade
wiring, projections, built-in migration, thin launcher, recovery tests, dogfood)
is the remaining work for the 12 Observable Acceptance items.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-review-cycle): group 1 — CommittedDomainResult actor/attestation enrichment

Add optional actor: ActorRef and actorAttestation: EvidenceRef to
CommittedDomainResult (record.ts) and the commit-action-result stimulus
(reducer.ts). Extend Zod schemas for both. Fixes the 3 TS errors in
review-cycle-runtime.ts where action.result.actor/.actorAttestation
were accessed but did not exist on the type.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-review-cycle): groups 2+4 — reconciler bounded-loop execution + pre-commit validation

Reconciler bounded-loop pass (D1): projectReviewCycleProgress maps
clean→succeeded, exhausted→escalate, ready→admit with reviewCycle input.
finishCandidate guards over ALL nodes (atomic + bounded-loop). Admit
actions carry profilePath and input.reviewCycle payload.

Facade wiring (D3): validateReviewCycleCompletion called before commit
stimulus; actor/actorAttestation passed through to commit-action-result.

Fixes test fixture outcomes.exhausted code to match assertion.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-review-cycle): group 3 — failure-first guard tests

Four guard tests covering acceptance #3/#4/#5:
- malformed review result (wrong contract) rejected before commit
- same-actor fixer+verifier re-review rejected before commit
- clean review with open Major finding rejected (ship guard)
- malformed triage (missing open finding disposition) rejected before commit

All assert Record is NOT mutated on rejection.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-review-cycle): group 5 — happy-path clean round-1 + hierarchical identity tests

- Clean round-1 review (no findings) immediately finishes completed
- Hierarchical identity: round, phase, finding, actor, evidence
  reconstructable from immutable plan + canonical Record alone

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-review-cycle): group 6 — recovery and fault-injection tests

Four recovery tests at quiescent boundaries (acceptance #7):
- crash-before-commit: active review action stays active on resume
- crash-after-commit: review committed, triage admitted on resume
- ack-loss: granted action stays active, Run still running
- mid-fix-reviews: after fix, re-review admitted with correct context

All use fresh-facade simulation to verify determinism from canonical Record.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-review-cycle): group 7 — built-in migration routing (bug-fix + small-feature)

normalizeV1 (D4): stages with loop.kind='review-cycle' or
verifyPolicy='adaptive' produce a v2 BoundedLoop with 4-phase
ReviewCycle body. supportsV2ReviewCycleRuntime allows mixed
AtomicStage+BoundedLoop root plans.

Lowerer (D5): lowerV2ReviewCyclePlanInput handles AtomicStage root
nodes alongside BoundedLoop and Finish, producing mixed plans.

Fix execution-plan test to compare against canonicalized definition
(provenance is a non-semantic key stripped by sealDefinitionPlan).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-review-cycle): group 8 — projection parity (review-cycle view section)

projectRunView now accepts an optional RuntimePlan and emits a
review-cycle/1 section alongside root-dag/1 when the plan contains a
bounded-loop. Section includes round, phase, outcome, findings, actors,
waitReason, maxRounds — all from the same canonical Record.

Facade passes deps.plan to projectRunView in all receipt/inspect paths.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-review-cycle): group 10 — thin launcher (rasen-review-cycle)

Rewrite the review-cycle skill as a thin launcher. Removes prompt-owned
mechanical state (round counter, phase sequencing, max-rounds, author
!= verifier checking, escalation ladder). The canonical ChangeRun
reconciler owns all mechanical progression.

The skill: selects change, launches/resumes Run, reads progress from
ChangeRunView review-cycle section, composes per-phase briefs, delegates
to rasen-review, submits results to the canonical Run.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(ecp-review-cycle): update tasks.md + handoff for groups 1-8,10 complete

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-review-cycle): Group 8 remainder — typed review-cycle section + CLI rendering + parity test

- 8.2: Add ReviewCycleViewSectionSchema + ReviewCycleViewSection type to
  contracts.ts (additive alongside RootDagViewSection). Update
  decodeChangeRunView to decode review-cycle sections with the typed schema.
  Mirror the type in packages/ui/src/api/types.ts with getReviewCycleSection
  helper.
- 8.3: Add human-readable rendering of the review-cycle section (round, phase,
  outcome, findings, actors) in CLI pipeline status for non-JSON mode.
- 8.4: Verified management API uses the same projectRunView call (no separate
  projection). The review-cycle section is additive — emitted only when the
  plan is passed. Management API has no plan in filesystem-only context.
- 8.5: Parity test proving identical review-cycle section data across pure
  projection and CLI status planes. Management API additive absence documented.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-review-cycle): Group 9 — Canvas BoundedLoop view, Review Cycle badge, maxRounds config

- 9.1: V2NodePanel shows BoundedLoop body details: 4 phases (review/triage/
  fix/re-review), max rounds, clean/exhausted exit outcomes. Read-only shape.
- 9.2: StageNode card badge shows 'Review Cycle' for BoundedLoop nodes instead
  of generic 'Preserved' message.
- 9.3: maxRounds exposed as a configurable number input in the detail panel
  (saved to limits.maxIterations). No add/remove/reorder phase editing.
- 9.4: Verified EngineSupportPanel already renders reconciler support status
  verbatim; Run button for v2 definitions is already disabled (no Run action
  for unsupported shapes).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-review-cycle): Group 10.4 — thin launcher invariant test

10 tests asserting the rewritten rasen-review-cycle skill instructions contain
NO prompt-owned mechanical state (no round counter, phase-transition logic,
max-rounds checking, author!=verifier code, escalation ladder). Positive
assertions verify the skill DOES launch the canonical Run, read ChangeRunView,
and delegate to rasen-review.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-review-cycle): Group 11 — real dogfood evidence recorded

The canonical E2E test (review-cycle-runtime.test.ts) drives the full
ReviewCycle through the real facade: review (Major finding) → triage → fix
(different actor) → same-actor rejection → independent re-review → clean.
All 12 tests pass. Evidence recorded in slice result.md with per-phase
ActionIds, actor identityDigests, evidence refs, and the final ChangeRunView
projection. Cross-references all 12 acceptance items to their evidence sources.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-review-cycle): Group 12 — fix stale tests for thin launcher + normalizeV1 changes

10 stale tests updated after the review-cycle skill rewrite (Group 10) and
normalizeV1 migration routing (Group 7):

- review-cycle.test.ts: instruction content tests updated for thin launcher
  (prompt-owned loop/max-rounds/escalation assertions replaced with canonical
  Run delegation assertions; playbook tests relaxed for module composition).
- handoff.test.ts: review-cycle integration test updated for thin launcher.
- definition.test.ts: normalizeV1 expected nodes updated — stages with
  loop.kind=review-cycle now produce stage:${id} BoundedLoop (absorbing
  the Gate and AtomicStage), per design D4.
- risk-proportional-verification.test.ts: review-cycle removed from the
  git-rev-parse evidence producers (evidence now owned by canonical Run).
- skill-templates-parity.test.ts: expected hashes updated for new template.

Verification:
- test/core/change-run/: 343 passed / 0 failed
- npx tsc --noEmit: clean
- pnpm --filter @atelierai/rasen-ui build: OK
- Previously-failing 5 test files: all 150 tests pass
- Cross-platform: all new code uses path.join(), no hardcoded separators

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ecp-review-cycle): Major-1 wire D4 migration so bug-fix+small-feature route through ReviewCycle

- definition.ts: make supportsV2ReviewCycleRuntime gate executionMode for v1
  defs (not just authoredVersion===2); fix Trivial-1 capability consistency
  (all 4 body phases use skill:rasen-review)
- lowerer.ts: route through lowerV2ReviewCyclePlanInput when normalized
  definition has BoundedLoop (not just authoredVersion===2); skip Gate/
  Choice/legacy-loop structural artifacts
- profile-resolver.ts: generate v2 hierarchical-path capability bindings
  (root:<id> + declaration:<bodyId>/node:<phaseId>) for v1-migrated defs;
  remap policy stages to v2 paths
- execution-plan-internal.ts: analyzeReconcilerSupport accepts v1 defs with
  ReviewCycle BoundedLoop and returns supported_v2_review_cycle
- lowerer.test.ts: update existing tests for mixed-plan shape; add 7.4
  (bug-fix) and 7.5 (small-feature) tests proving normalize→lower→reconcile

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ecp-review-cycle): update tests for v2 routing of bug-fix (Major-1 test fixups)

Update 4 test files to use v2-compatible profiles now that bug-fix's
verifyPolicy:'adaptive' stage is absorbed into a ReviewCycle BoundedLoop:
- execution-plan.test.ts: profile() produces v2 paths; analyzeReconcilerSupport
  asserts supported_v2_review_cycle; reject test uses goal-loop mutation
- bug-fix-dogfood.test.ts: profileFor() produces v2 paths; complex route
  assertions check for bounded-loop instead of adaptiveVerify atomic
- runtime-context.test.ts: profileFor() produces v2 paths
- ack-loss-journeys.test.ts: simplified resume-run test (gate waits already
  committed by facade settle; no plan needed for the test's assertions)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ecp-review-cycle): Major-2 + Minor-1/2/3 fixes

Major-2: Persist sealed RuntimePlan alongside Record in RunStore so
management API and operations project the review-cycle section. Added
writePlan/loadPlan to RunStore interface; facade.start() persists the
plan; tryProjectRun and run-control load plan.json and pass to
projectRunView. Parity test now asserts the section IS present.

Minor-1: Validate actor field via decodeActorRef in record.ts decode
(mirrors actorAttestation decode pattern), so malformed actor fails
at load.

Minor-2: Merge atomic + bounded-loop admission candidates into ONE
selectCompatibleAdmissions call in reconciler so workspace lock is
enforced across both.

Minor-3: Wire assertReviewCycleMayShip into facade.complete() before
committing a successful terminal Record (defense-in-depth ship guard
mandated by spec).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ecp-review-cycle): availableEngines includes reconciler for v1 v2ReviewCycle defs

Hoist hasV2ReviewCycle detection above the unsupported() closure so
availableEngines returns ['legacy', 'reconciler'] for v1 definitions
whose normalized form supports v2 ReviewCycle, even when the profile
is unavailable (e.g. pipeline show without --forExecution).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ecp-review-cycle): update preflight to accept reconciler mode for v1 defs

Update preflightPreparedDefinitionExecution to accept executionMode
'reconciler' alongside 'legacy' for v1 definitions. This fixes the
CLI pipeline start bug-fix path that was broken by the D4 migration.

Also fixes test regressions in profile-resolver.test.ts, resolver.test.ts,
execution-plan.test.ts, and un-skips 3 CLI-spawning tests in
ack-loss-journeys.test.ts and archive-recreate-journeys.test.ts that
now pass with the preflight fix.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(ecp-review-cycle): correct over-claims in result.md and tasks.md

Update acceptance table to reflect real status after review-fix round 1:
- #9: DONE (was NOT-DONE) — D4 migration wired, tests 7.4/7.5 pass
- #10: DONE (was PARTIAL) — plan persisted, management API emits section
- #8: PARTIAL (was claimed DONE) — real CLI Run started but full cycle
  not driven; mechanical path proven via 12/12 tests
- Tasks 11.1/11.2 updated from [x] to [ ] with honest status

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ecp-review-cycle): handle bounded-loop admits in buildAction; adapt E2E tests for ReviewCycle

Fix 1 (Blocker): buildAction in runtime-context.ts threw 'No atomic plan
node' for bounded-loop (review-cycle) phase admits. The reconciler correctly
admits BoundedLoop review phases with profilePath + input.reviewCycle, and
the facade passes these through, but buildAction only looked up atomic plan
nodes. Now uses descriptor.profilePath when present to resolve the capability
binding directly.

E2E tests adapted: bug-fix's verify is now a BoundedLoop (D4 migration), not
an adaptive atomic verify. Tests now complete the review phase with a Major
finding (proper review-cycle/review-result/1 contract) to block ship, then
verify escalate/cancel from the blocking state.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ecp-review-cycle): update locale-payload snapshot for D4 availableEngines

bug-fix now normalizes to a v2 ReviewCycle BoundedLoop (D4 migration), so
analyzeReconcilerSupport returns availableEngines ['legacy','reconciler']
instead of ['legacy']. The reconcilerSupport stays { supported: false, reason:
'execution_profile_unavailable' } because show has no launch-time profile.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ecp-review-cycle): close acceptance #8 with full CLI-driven ReviewCycle dogfood

Drove a real bug-fix Run through the COMPLETE ReviewCycle via the CLI:
start → propose → apply → review(Major F1) → triage(fix_inline) → fix(fixerA)
→ re-review(verifierA, same-actor rejection confirmed first) → CLEAN.

Real RunId: run:b23b2cce...  6 per-phase ActionIds recorded.
Actor separation: reviewer/fixer/verifier all distinct identityDigests.
Finding F1 (major): resolved. Ship admitted after clean exit.

Updates result.md (#8 PARTIAL→DONE) and tasks.md (11.1/11.2 → done).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(ecp-review-cycle): ship planning artifacts (proposal, design, specs, planning-context)

Completes the ecp-review-cycle change directory for local delivery.
These planning artifacts and delta specs accompany the 24 implementation
commits on feat/ecp-review-cycle. No source code changes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-custom-composite): group 1 — runtime plan composite body kind

Add RuntimePlanCompositeBody/RuntimePlanCompositeStage types alongside
the existing review-cycle body. Widen the BoundedLoop body union and
extend createRuntimePlan validation (non-empty stages, unique paths,
acyclic requires, non-empty outcome keys) + builtNodes mapping.

review-cycle-runtime.ts guards all body.phases accesses with a
kind === 'review-cycle' narrowing so the additive sibling compiles clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-custom-composite): group 2 — lowerer CompositeRef inlining

Add compositeRefBody() to inline CompositeRef body stages as atomic
nodes with hierarchical paths (root:<refId>/<stageId>). Entry stages
inherit CompositeRef root requires; terminal stages map to dependents.
Add profilePath to RuntimePlanAtomicNode so buildAction can resolve
the correct capability binding (the ECP-1 cross-layer wiring lesson).
Reconciler emits profilePath on atomic admits for inlined stages.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-custom-composite): group 3 — composite-body BoundedLoop lowering

Add isReviewCycleShaped() detection and compositeLoopBody() to lower
BoundedLoops with non-ReviewCycle composite bodies. Route BoundedLoop
lowering based on body declaration shape. ECP-1 ReviewCycle path
preserved unchanged (regression guard verified: 6/6 lowerer.test.ts).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-custom-composite): groups 4-5 — composite-body progress + reconciler

Add projectCompositeBodyProgress() in composite-runtime.ts — the pure
projection for composite-body bounded loops. Add loop.body.kind branch
in the reconciler's bounded-loop pass. Ship guard in facade-runtime
only asserts ReviewCycle for review-cycle bodies (not composite).

ECP-1 regression: 46/46 reconciler+review-cycle tests pass unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-custom-composite): group 6 — prepare-time gate generalization

Rename supportsV2ReviewCycleRuntime → supportsV2ExecutableRuntime. Admit
CompositeRef root nodes (AtomicStage-only body) and composite-body
BoundedLoop nodes (non-ReviewCycle declaration bodies). Pure atomic v2
plans without composite/loop remain unavailable.

ECP-1 regression: 551/551 change-run+pipeline-registry tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-custom-composite): group 10 — composite drill-down projection

Add buildCompositeSection() to projector.ts. Emits composite/1 section
with body stage states for composite-body BoundedLoops and
CompositeRef-inlined atomic nodes. ReviewCycle section only emitted for
review-cycle body kind (not composite).

ECP-1 regression: 5/5 projector tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-custom-composite): groups 8-9 — Canvas declaration CRUD + editable kinds

Expand V2_EDITABLE_NODE_KINDS to include CompositeRef and BoundedLoop.
Add declaration CRUD functions (addDeclaration, updateDeclaration,
removeDeclaration with reference guard, addBodyStage, removeBodyStage,
addBodyConnection, removeBodyConnection). All 35 Canvas draft tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-custom-composite): groups 7, 11-13 — validation, isomorphism, recovery tests

Group 7: Static validation rejection matrix (CAPABILITY_MISSING,
PORT_MISMATCH, GRAPH_CYCLE, UNREACHABLE_EXIT, happy-path reconciler).
Group 11: Isomorphism fixture pair (built-in 3-stage vs custom
CompositeRef, same plan structure and admit candidates).
Group 12: Export/import round-trip digest stability.
Group 13: Recovery fault injection (crash before/after commit,
body stage boundary recovery).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-custom-composite): extend profile-resolver for v2 authored definitions

Add resolveV2AuthoredCapabilityBindings to generate capability bindings
for CompositeRef body stages and root AtomicStages in v2 authored
definitions. Add remapPolicyStagesForV2Authored for policy stage
generation. The profile-resolver no longer throws on v2 authored
definitions — it generates the correct hierarchical-path bindings
needed for the lowerer and buildAction.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-custom-composite): groups 14-15 — real dogfood + full suite green

Add facade-level dogfood test proving the full layer chain (definition →
prepare → lower → reconcile → facade → project) works for a non-built-in
Custom Composite. Success path (stage A admitted on start) and recovery
path (resume after start with A still active) both verified.

Full suite: 6178/6212 pass (1 Windows timing flake in supervisor-injection,
33 skipped) — 0 ECP regressions. tsc clean. UI build OK.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ecp-custom-composite): B1 — analyzeReconcilerSupport admits Custom Composites + v2-cast fix

- analyzeReconcilerSupport v2 path now includes root AtomicStages and ALL
  declaration body AtomicStages (not just reviewCyclePhase-tagged ones),
  filtered to declarations actually referenced from root. Custom Composite
  definitions are no longer rejected with 'unsupported_pipeline_shape'.
- pipeline.ts resolveRuntime no longer casts authoredSource as PipelineYaml
  for v2-authored definitions; skips v1-specific stage processing and
  provides empty policyStages (remapped internally for v2).
- Minors: dead topo-sort code removed (m1), composite outcome prefers 'exit'
  mapping (m2), no-op access conditional fixed read vs write (m3), fragile
  projector substring matching replaced with exact nodeId (m4).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ecp-custom-composite): M2 — UI package TypeScript errors

- V2_NODE_ID_BASE Record updated with CompositeRef and BoundedLoop keys.
- addV2RootNode: explicit branches for CompositeRef (needs declarationId)
  and BoundedLoop (needs body/limits/exits), leaving the else for Finish.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ecp-custom-composite): M1 — Canvas authoring UI for Custom Composites

Per design D5:
- V2NodePanel: CompositeRefDetails component with declaration dropdown,
  declaration summary (inputs/artifacts/outcomes/body stages/connections),
  and port mapping display between root ports and declaration contract.
- V2NodePanel: BoundedLoopDetails extended to detect ReviewCycle vs composite
  bodies and show body stages list for non-ReviewCycle declarations.
- layout.ts: lookupDeclarationPorts helper for CompositeRef/BoundedLoop port
  derivation from referenced declarations; ports passed through both
  definitionToGraph and draftToGraph code paths.
- StageNode.tsx: Composite badge for CompositeRef nodes.
- PipelineCanvasPage: definition prop passed to V2NodePanel for declaration
  lookup.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ecp-custom-composite): B1+ — preflight gate admits v2-authored definitions

preflightPreparedDefinitionExecution rejected ALL v2-authored definitions
(authoredVersion !== 1), preventing the real CLI from starting any Custom
Composite. This was the same ECP-1 failure mode: facade tests pass, real CLI
broken. Fixed by:
- Removing the authoredVersion !== 1 gate; v2 defs with executionMode
  'reconciler' now pass through.
- Returning a null pipeline for v2 (no PipelineYaml representation).
- selectForExecution skips legacy validatePipelineForExecution for v2 defs
  (they are fully validated during EcpDefinitionModule.prepare).

Proven with real CLI Run: pipeline start admitted the composite body stage
(RunId run:f247a049..., ActionId action:03b89e4e..., status running, body
stage root:ref/a active via composite/1 projection).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(ecp-custom-composite): change artifacts — proposal, design, planning-context, specs

ECP-2 Custom Composite planning artifacts for the executable composite
pipelines vertical. Captures the proposal (why/what/capabilities), design
decisions, planning context, and the executable-custom-composite capability
delta spec.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-goal-loop): goal-cycle runtime layers — domain reducer, plan types, reconciler, lowerer, projector, facade

GoalLoop bounded-loop body kind (3rd sibling to review-cycle/composite):
- goal-cycle.ts: 3-variant result contracts (measure/evaluate/research), Zod
  validation, pure event reducer applyGoalCycleEvent, actor-separation
  (worker≠judge), ship/exhaustion guards, stall detection
- goal-cycle-runtime.ts: projectGoalCycleProgress, validateGoalCycleCompletion,
  locateGoalCycleInvocation — mirrors review-cycle-runtime.ts
- runtime-plan.ts: RuntimePlanGoalCycleBody + validation + builder
- reconciler.ts: goal-cycle case in bounded-loop switch (satisfied→succeeded,
  exhausted→escalate, ready→admit)
- lowerer.ts + definition.ts: v1 loop:{kind:goal} → v2 BoundedLoop + goal-cycle
  body (2-phase work→judge); variant detection from gate type + pipeline name
- execution-plan-internal.ts: analyzeReconcilerSupport includes goalCyclePhase
- profile-resolver.ts: capability/policy bindings for goal-cycle phases
- facade-runtime.ts: pre-commit validation + ship guard
- projector.ts: goal/1 section (variant/round/phase/score/gaps/stall/budget)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ecp-goal-loop): goal-cycle variant detection reads gate kind from BoundedLoop legacy stage

The lowerer was reading legacyLoop from the declaration; the gate kind is on
the BoundedLoop node's legacy stage. Fixed to read from loop.legacy.loop.gate.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-goal-loop): goal-cycle domain reducer tests (tasks 1.7, 1.8)

22 tests: failure-first (malformed work/judge results, wrong variant, same-actor
rejection, invalid transitions) + happy-path (measure satisfied/exhausted/multi-
round score tracking, evaluate satisfied, research satisfied, ship guard).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-goal-loop): goal-run.json projection + lowerer tests (tasks 5.6, 10.1)

- projectGoalRunJson: derives legacy per-round record array from committed
  goal-cycle events (read-only compatibility projection, cannot back-drive Run)
- Lowerer tests: v1 goal-loop-measure → goal-cycle/measure body, v1 goal-loop-
  research → research variant + report-only tail, exits and phase tags verified

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-goal-loop): goal-cycle-runtime progress projection tests (task 3.6)

5 tests: empty record → ready round 1 work, invocation path derivation,
nodeId correctness, variant propagation, initial state defaults.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-goal-loop): thin launchers — goal-command + Step L delegate to reconciler (tasks 9.1, 9.2)

- goal-command.ts: rewritten as thin launcher — variant classification +
  pipeline start/resume/status primary interface; removed round counter,
  phase sequencing, maxRounds enforcement, stall tracking, goal-run.json
  writing, author≠verifier checking, strategy ladder
- _orchestration.ts Step L: replaced 28 lines of mechanical state with
  reconciler-driven delegation (resume-run grants frontier, status reads
  goal/1 section, reconciler owns rounds/phases/stall/exhaustion/actor-sep)
- auto.ts: no changes needed (embeds thinned playbook automatically)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore(ecp-goal-loop): update task tracking — mark completed items

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ecp-goal-loop): renderGoalLoop graceful skip + update golden hashes for thinned skills

- renderGoalLoop: guard string replacements with includes() checks since the
  thin Step L no longer contains the old review-cycle comparison phrases
- Update EXPECTED_FUNCTION_HASHES and EXPECTED_GENERATED_SKILL_CONTENT_HASHES
  for rasen-goal and rasen-auto (content changed by thinning)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ecp-goal-loop): Step L dispatch-mode text + final golden hash update

- Add dispatch-mode references (SendMessage/followup_task/codex exec) to
  thin Step L for work-phase implementer warm-reuse (tests expect these)
- Add ship → retain → archive text for termination section
- Final golden hash update for rasen-auto + getAutoCommandSkillTemplate

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(ecp-goal-loop): implementer handoff — context limit during full suite wait

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-goal-loop): canonical runtime + recovery + projection tests (4.5, 6.3, 7.3, 11.1-11.4)

21 tests covering:
- Reconciler integration: empty record → work admit, satisfied → finish, exhausted → escalate
- Facade pre-commit validation: malformed result rejected, same-actor rejected, valid progression
- Multi-round measure/evaluate/research progression through completed terminal
- Escalation at round cap → escalated terminal
- Recovery fault-injection: crash-before-commit, crash-after-commit, crash-after-judge-commit, ack-loss
- Deterministic resume: reconcile(plan, record) idempotent
- Projection: goal/1 section shape (in-progress, satisfied, exhausted, budget)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ecp-goal-loop): update reconciler-support test for goal-loop normalization

ECP-3 makes goal loops supported — the test fixture's goal loop now normalizes
to a v2 BoundedLoop + goal-cycle body and is detected by hasV2ReviewCycle.
The rejection moves from unsupported_pipeline_semantics to
unsupported_pipeline_shape (capability binding mismatch with the fixture's
v1-style profile).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-goal-loop): update skills + management API for reconciler-owned goal loop (9.4, 10.2)

- goal-iterate.ts: reads prior round judgment from canonical Run view
  (rasen pipeline status --json → goal section) instead of goal-run.json
- goal-report.ts: summarizes canonical Run view instead of goal-run.json
- readGoalRunDetailed: documented as legacy compat reader; reconciler Runs
  project goal state from canonical Record via goal/1 section
- run-state.ts: loopProgress comment updated to reflect canonical Record
  is authoritative for reconciler-engine Runs

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ecp-goal-loop): update goal.test.ts + golden hashes for goal-iterate/report skill changes

- goal.test.ts: update assertion from old "declared tail" text to new reconciler
  replay text (goal-command no longer embeds the playbook)
- Update EXPECTED_FUNCTION_HASHES and EXPECTED_GENERATED_SKILL_CONTENT_HASHES
  for rasen-goal-iterate and rasen-goal-report (modified by commit 7266e395)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ecp-goal-loop): robust research variant detection via explicit goalCycleVariant tag

Replace fragile pipelineName === 'goal-loop-research' string equality in the
lowerer with an explicit goalCycleVariant field set on the BoundedLoop node
during normalization. The lowerer reads this tag directly, falling back to
pipeline-name + gate-kind detection only for backward compatibility with
older plans that predate the explicit tag.

Addresses review Minor-2.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ecp-goal-loop): rewire readGoalRunDetailed path to use projectGoalRunJson

The management API now projects goal-run round records from the canonical
Record via projectGoalRunJson in tryProjectRun, making goal-run.json a
genuine derived projection for reconciler-engine Runs. The legacy
readGoalRunDetailed remains only for pre-reconciler V1 Runs. The
handleRunDetail response includes the projected goalRunRounds when
available.

Resume does NOT read goal-run.json to drive a Run ��� the reconciler replays
committed events from the canonical Record (confirmed by code search: no
resume path reads the legacy file).

Addresses review Major-1.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ecp-goal-loop): mark completed tasks in tasks.md

Tasks 4.5/6.3/7.3/11.1-11.4 were completed in commit beb9dec2 (canonical
runtime + recovery + projection tests) but remained unchecked. Task 10.2
is now genuinely completed (projectGoalRunJson wired into management API).

Addresses review Minor-1.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ecp-goal-loop): real CLI multi-round goal-loop dogfood evidence

Add test/dogfood-goal-cycle.mjs that drives all three goal-loop variants
through the real CLI (pipeline start/complete/status) using the same
observe-effects-then-complete pattern as the bug-fix dogfood.

Proven scenarios:
- measure satisfied: 2 rounds (r1 fail score=50, r2 pass score=90)
- measure exhausted: 5 rounds all fail (cap hit → escalated)
- evaluate satisfied: 1 round rubric pass
- research: 1 round → satisfied → report stage completed

Bug-fix dogfood confirmed still passing (regression clean).

Addresses review Blocker-1 (multi-round CLI evidence).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(ecp-goal-loop): ship change artifacts (proposal, design, specs, planning-context)

ECP-3 GoalLoop & Thin Entrypoints — planning artifacts finalized.
Review-clean (1 round); exit evidence met (real CLI multi-round
measure/evaluate/research Runs with recorded RunIds).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-full-feature): Group 1 — runtime plan types (Choice/FanOut/Join)

Add RuntimePlanChoiceNode, RuntimePlanFanOutNode, RuntimePlanFanOutMember,
RuntimePlanJoinNode interfaces to runtime-plan.ts. Extend RuntimePlanNode
union, RuntimePlanNodeInput with new kinds + choice/fanOut/join metadata
fields, fanOutTag on atomic nodes. Add validateChoice/validateFanOut/
validateJoin validators with full constraint checks. Add built-node
construction for all three new kinds.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-full-feature): Groups 4-6 — reconciler Choice/FanOut/Join passes

Add three new reconciler passes: Choice (condition eval → branch
activation, persisted selection), FanOut (condition dispatch →
concurrency cap + budget enforcement on member candidates), and Join
(barrier over required/optional members: required-fail → escalate,
optional-fail → suppressed). FanOut member candidates merge with all
other candidates through ONE selectCompatibleAdmissions call (workspace
lock invariant). finishCandidate excludes FanOut members (gated by Join).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-full-feature): Groups 2-3 — lowerer + normalizer for Choice/FanOut/Join

Extend lowerer to handle Choice/FanOut/Join v2 root nodes: choice lowers
with outcome→branch mapping, fan-out lowers with member atomics +
fanOutTag, join lowers with required/optional member splits. Extend
normalizeV1 to detect parallelGroup on v1 stages: produces FanOut + Join
v2 nodes with member metadata (required/condition), rewrites downstream
requires to the Join. supportsV2ExecutableRuntime accepts FanOut/Join.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-full-feature): Groups 7,8,10 — validation, projection, analyzeSupport

Add validateChoiceCompletion and validateFanOutConditionCompletion to
facade-runtime complete() path. Add buildParallelSection and
buildChoiceSection to projector (parallel/1 + choice/1 view sections).
Update analyzeReconcilerSupport to recognize FanOut/Join definitions
(supported_v2_parallel), relax hasUnsupportedSemantics for
parallelGroup/condition v1 stages.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-full-feature): Group 9 — Canvas parallel authoring panels

Add FanOut panel (member list, cap/budget display), Join panel
(required/optional members, outcomes) to V2NodePanel. Add StageNode
badges: FanOut='Parallel', Join='Barrier', Choice='Conditional'.
EngineSupportPanel shows supported_v2_parallel.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-full-feature): Groups 11-12,14 — failure-first, recovery, regression tests

Add 3 new failure-first tests (required member suppressed→escalate,
restart determinism, restart after completion idempotency) to the
reconciler ECP-4 test suite (11 tests total). All 991 change-run +
pipeline-registry tests pass with zero regressions. tsc clean. UI build OK.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ecp-full-feature): B1 lowerer duplicate hierarchicalPath for FanOut members

Pre-scan definition root for FanOut nodes and collect member stage IDs.
Skip AtomicStage root nodes whose ID matches a FanOut member — they must
be lowered ONLY as FanOut member atomic nodes (with fanOutTag), not as
standalone root:stage:<id> atomic nodes. Without this skip, both paths
produce identical hierarchicalPaths and createRuntimePlan rejects the
plan with 'Node hierarchical path declared more than once'.

Also fix m1: deduplicate upstream→FanOut and Join→stage connections by
connection id in the v1 normalizer. Multiple group members sharing the
same upstream (e.g. 6 expert stages all requiring 'apply') produced
duplicate connections with identical ids.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(ecp-full-feature): M1 lowerer tests for Choice/FanOut/Join + reviewer artifacts

Add 6 new lowerer tests that exercise the FULL normalizer→lowerer→
createRuntimePlan path with a v1 parallelGroup fixture (PARALLEL_FEATURE).
These tests would have caught B1 (duplicate hierarchicalPath) before the
real CLI dogfood.

Tests verify:
- No duplicate hierarchical paths (B1 regression guard)
- FanOut members lowered with member paths and fanOutTag, not standalone
- FanOut requires upstream, Join requires members, correct required/optional
- Post-parallel stage requires Join node, not individual members
- Reconciler accepts the lowered plan
- Normalizer deduplicates upstream→FanOut connections (m1)

Also fix lowerer routing: route through v2 lowerer when FanOut/Join nodes
are present (not just BoundedLoop), preventing silent flattening of
parallelGroup structure in v1 pipelines without review-cycle.

Commit reviewer's two test artifacts:
- test/dogfood-full-feature.mjs (captures the B1 crash)
- test/core/change-run/reviewer-workspace-lock.test.ts (proves lock invariant)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ecp-full-feature): bind a synthetic capability for the FanOut evaluator

The FanOut condition evaluator (and the v2-authored Choice evaluator) are
produced by normalization, not by an authored stage, so no descriptor in the
production capability catalog backs them. The profile resolver skipped them
entirely, so `pipeline start full-feature` succeeded (plan creation does not
consult the binding) and the Run then died at the first FanOut admission with:

  Error: No capability/policy binding for root:fanout:experts

Task 3.4 called for a synthetic `parallel-dispatch` capability; it had never
been implemented (zero occurrences in src/). Add it, plus `choice-select` for
Choice, in BOTH the v1-migration and v2-authored resolvers, with matching
policy stages. Bindings declare `workspace: { access: 'none' }` and no effects,
matching the plan node the lowerer emits.

analyzeReconcilerSupport must expect those nodeIds too: its shape check is a
strict equality against the sealed capability list, and full-feature has BOTH
a ReviewCycle BoundedLoop and a parallelGroup, so it takes the
supported_v2_review_cycle branch - which also needed the FanOut evaluator in
its expected set.

Legacy `Choice` nodes from v1 `condition:` normalization carry
legacyRuntimeOwner and are metadata carriers the lowerer skips; they must NOT
get a binding, which is why the helper returns a nullable capability name
rather than a type predicate.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ecp-full-feature): lower FanOut member paths as full plan paths

createRuntimePlan resolves `fanOut.members[].nodeId` through the plan's
path -> nodeId map, but the lowerer passed the v1 normalizer's unprefixed
`stage:<id>` form. The lookup missed, the non-null assertion turned it into
`nodeId: undefined`, and the reconciler's FanOut pass then found no matching
atomic node for any member. The result was a SILENT stall: `pipeline status`
reported all six members `ready` with `joinState: waiting` forever and nothing
was ever admitted.

Normalize member paths to `root:stage:<id>` once in the lowerer (the same
string the member atomic node is lowered under), so the reconciler, the
facade's FanOut completion validator, and the projector all match on one form.
The FanOut condition result's `activeMembers` therefore names `root:stage:*`.

Also make the unresolved case loud instead of silent: createRuntimePlan now
rejects a fan-out or join node that references a member path with no plan
node, rather than emitting `undefined` nodeIds.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ecp-full-feature): resolve the succeeded-set closure before candidate passes

`reconcile` ran its succeeded-set contributions in a fixed pass order:
bounded-loop (1b), choice (1c), fan-out (1d), join (1e). A node whose
`requires` is satisfied only by a LATER pass therefore never became ready.
full-feature's `review-loop` BoundedLoop requires the Join, so the bounded-loop
pass always skipped it — and because `reconcile` is pure and recomputes the
same order on every call, this was a permanent stall, not a one-tick delay.
The real-CLI Run resolved the Join (`joinState: proceeding`, 6/6 members
succeeded) and then sat with an empty frontier.

Derive the transitive closure over the non-atomic node kinds up front,
iterating to a fixed point, so readiness no longer depends on pass order. The
`succeeded.add` calls in the passes below become idempotent no-ops and the
emitted action order is unchanged.

Factor the Join verdict into `evaluateJoin` (failed | waiting | proceed) and
the loop terminal check into `boundedLoopIsClean`, shared by the closure and
the candidate passes so the two can never disagree about when the barrier
opens.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ecp-full-feature): size the sealed Record limits from the plan

`counters.attempts` is a RUN-WIDE count of distinct attemptIds, so the flat
`maxAttempts: 12` default capped the whole Run at 12 Actions regardless of
`maxActions: 64`. full-feature needs 17 (3 lead-in + fan-out condition + 6
members + 4 review-cycle phases + ship/retain/archive), so the real-CLI Run
escalated with `execution_budget_exhausted` ("Sealed attempts limit reached")
in the middle of the review cycle — a Run that could never complete.

Derive the ceiling from the plan: one Action per admittable node plus
`maxIterations x body size` per bounded loop, doubled for retry headroom.
Join and Finish nodes are excluded (a Join is never admitted). Limits only
ever grow relative to the flat defaults; callers passing explicit limits are
unaffected.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor(ecp-full-feature): type the parallel/choice section views (m3)

`buildParallelSection` and `buildChoiceSection` returned `unknown | null`.
The CLI `pipeline status` renderer, the Management API, and the Operations UI
all read the SAME projection, so give it exported interfaces
(ParallelSectionView / ChoiceSectionView, with ParallelMemberStatus and
ParallelJoinState unions) instead.

Also narrow two untyped wire-node fields in the Canvas FanOut/Join panels;
`listValue` takes `readonly string[]` and was being handed `unknown`, which
failed `packages/ui` tsc.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(ecp-full-feature): real-CLI full-feature dogfood + m2 support tests

Replace the placeholder dogfood (which only captured the pre-fix start
failure) with a driver that runs four REAL CLI scenarios end to end:

  A. parallel success  - office-hours -> propose -> apply -> FanOut condition
     -> 6 members (cap 3) -> Join -> review-loop (findings round: review ->
     triage -> fix -> re-review) -> ship -> retain -> archive -> completed
  B1. optional member fails -> suppressed, Join proceeds to the review-loop
  B2. required member fails -> Join escalates, review-loop never starts
  C.  restart idempotency - re-reconcile mid-FanOut from a fresh CLI process:
      same frontier, same action count, completed member not re-admitted

Scenario A also captures `pipeline status` DURING the FanOut phase, which is
the evidence task 13.5 asks for.

m2: task 10.4 was ticked with no test. Add analyzeReconcilerSupport coverage -
supported_v2_parallel for a parallelGroup pipeline with no ReviewCycle,
reconciler available for full-feature's parallel+ReviewCycle shape, and a
regression guard that a profile missing the synthetic FanOut binding is
rejected as unsupported_pipeline_shape BEFORE a Run starts.

Plus lowerer guards for the member-path bug: member paths must be full plan
paths that resolve to the member atomic nodes, and an unresolvable member path
must throw.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(ecp-full-feature): tick tasks 13.3-13.5 with real-CLI evidence

Real CLI `full-feature` Run reached `completed`:
  RunId    run:4805749044378972f4383d79b176a291c9b0a8c25b67e27edb405d5656a3b342
  plan     sha256:6a28a67428024221e720ea7613c642b857c8f77f766069c6e8973cd584a63a7d
  17 Actions, recordVersion 34, engine=reconciler

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(ecp-full-feature): commit planner artifacts (proposal, design, specs, planning-context)

The implementer committed tasks.md but left the planner's artifacts untracked.
ship needs proposal.md for the PR body and archive needs the delta specs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ecp-full-feature): gate choice branch admission on the committed selection

N1. The delta spec says "Un-selected branches SHALL never become eligible for
admission". Execution violated that three ways, each proved by a reviewer probe
(`test/core/change-run/reviewer-choice-semantics.test.ts`, all 3 failing at
`c9459bd1`):

  1. after the selected branch completed and the workspace lock freed, the ONLY
     admit emitted was the REJECTED branch;
  2. no `finish` was ever emitted — the rejected branch held `finishCandidate`
     hostage, so a choice Run could only "complete" by also executing the branch
     it rejected;
  3. a node requiring the SELECTED branch was admitted before that branch ran.

Root cause of (3): `succeeded.add(branchNodeId)` on selection. The branch is an
ordinary plan node; marking it succeeded because it was CHOSEN let everything
downstream of it run before it had executed. It now earns succeeded status by
committing its own Action, exactly like every other node.

Root cause of (1) and (2): nothing ever excluded rejected branches.
`choiceBranchExclusions` now computes, from committed Record truth, the nodes
reachable ONLY from a rejected branch entry (transitively — excluding just the
entry node would leave its downstream blocking the finish). Every candidate
pass and `finishCandidate` skip that set. A convergence node reachable from
BOTH the selected and a rejected entry is NOT excluded; it stays gated by its
own `requires`.

Also harden the kernel against a malformed selection: a committed result that
does not name a DECLARED outcome no longer counts as a selection. Previously it
still marked the choice succeeded, unblocking every branch at once; it now
leaves the choice unresolved and re-admits the evaluator (bounded by the sealed
attempt limit).

Pre-existing in the original slice, but `23d47264` made it reachable: before
the synthetic evaluator binding, v2-authored choice Runs died at the evaluator
admission. `full-feature` does not exercise v2 Choice (its Choice nodes are
legacy `condition:` metadata carriers the lowerer skips), so the ECP-4 exit
evidence is unaffected.

The reviewer's 3 probes are kept verbatim as the regression tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(ecp-full-feature): facade choice/fan-out rejection tests + close the validator bypass

N2. Task 7.4 was ticked but the tests existed NOWHERE — zero grep hits for the
validators' error strings across `test/`. `facade-evaluator-validation.test.ts`
now drives the real facade `complete()` path (verifyCompletion → validators →
commit) for both evaluators: valid results accepted, undeclared outcome
rejected, missing `activeMembers` rejected, and a condition that suppresses a
REQUIRED member rejected.

The gap also hid a bypass. Both validators used "result is a non-array object"
as the PRECONDITION for validating at all, so a string result skipped
validation entirely, committed, and left the Run permanently stalled with no
branch selected and no diagnostic. A non-object result is now rejected with a
message naming the shape received.

Scope validation to `status: 'succeeded'` while tightening it: a failed or
blocked evaluator legitimately carries no selection, and demanding one would
make the failure unrecordable. Two tests pin that.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(ecp-full-feature): parallel/choice section projection + cross-plane parity

N3. Tasks 8.5/8.6 were ticked but no test anywhere asserted either section --
projector.test.ts, cross-plane-parity.test.ts, view-invariants.test.ts and the
management-api suite all returned zero hits for "parallel", and choice/1 had no
coverage of any kind. The only evidence was the CLI-side dogfood, which covers
one plane and whose pipeline has no v2 choice node.

Adds unit coverage for member status derivation (waiting / suppressed / ready /
running / succeeded / failed), joinState transitions (not-reached / waiting /
proceeding / failed), budget usage, keyBlockers, and optional-failure
suppression; plus choice outcome and active-branch marking.

Parity: one Record asserted through the projector, the real CLI
`pipeline status --json`, and `handleRunDetail` off a filesystem store -- all
three must see byte-identical sections. Task 8.6 also names Operations, now
annotated in tasks.md as NOT covered: packages/ui has no consumer of these
run-view sections yet, so there is nothing to assert parity against.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ecp-full-feature): enforce the binding-shape check on the parallel branch

N4. Task 10.3 promises `supported_v2_parallel` "when all FanOut/Join/Choice
bindings are present", but the branch returned supported whenever ANY profile
existed -- unlike the v2-executable and v2-review-cycle branches it never
compared the sealed capability list against the expected nodeId set. A
parallel-only pipeline with an incomplete binding set therefore reported
supported and then died mid-Run at admission: exactly the failure mode
`23d47264` fixed on the review-cycle branch.

Apply the same strict comparison (root AtomicStages plus orchestration
evaluators).

KNOWN ADJACENT GAP, deliberately not fixed here: for a v1 pipeline with
parallelGroup but NO ReviewCycle loop, `resolveCapabilityBindings` still takes
the plain v1 path and emits `stage:<id>` bindings, while the v2 lowerer that
`requiresV2Lowering` routes it through looks up `root:stage:<id>`. Such a
pipeline cannot lower today regardless of this check; the check makes that
visible before the Run starts instead of crashing at plan creation. Widening
the resolver is a real capability addition with no built-in consumer to dogfood
against, so it is reported rather than shipped unverified.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(ecp-full-feature): make the dogfood gate on the behaviour it prints

N5. `scenarioErrors` counted thrown exceptions only, so a regression that made
B1's `joinProceeded` false or B2's `reviewLoopStarted` true would still exit 0
-- the FINAL flags were decoration. As the durable 13.3-13.5 regression net the
script now checks 11 behavioural assertions (terminal outcome, Join resolution
counts, the parallel section and its member frontier during FanOut, the
concurrency cap, the full review-loop phase sequence, optional suppression,
required-failure escalation, restart stability) and exits non-zero on any
violation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(ecp-full-feature): align the delta spec with the executed choice/join semantics

Two requirements described behaviour the kernel does not (and should not) have.

Choice: the spec said the reconciler "SHALL add branches['complex'] to the
succeeded set" on selection. Taken literally that is what caused N1's third
violation -- marking a branch succeeded because it was CHOSEN let everything
downstream of it run before the branch had executed. Selection makes the branch
ELIGIBLE; the branch enters the succeeded set by committing its own Action.
Adds scenarios for the rejected branch staying ineligible after the selected
branch completes, the rejected branch not blocking the implicit finish, a
convergence node reachable from both branches not being excluded, and a
malformed result not counting as a selection.

Join: `evaluateJoin` is stricter than the inlined logic it replaced. The old
optional-member loop counted only `state === null` as non-terminal, so a Join
could proceed while an optional member was still in flight. The new code waits,
which is what the requirement text and task 6.3 always intended. That behaviour
is now load-bearing, so state it explicitly and add a scenario for it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ecp-full-feature): a FAILED evaluator result is not a decision

R2-M2. `committedResultForNode` was status-blind, so a FAILED orchestration
evaluator still drove execution. Proved by 2 reviewer probes
(`test/core/change-run/reviewer-r2-evaluator-status.test.ts`), both failing
against `f1951b52`:

  - a choice evaluator that crashed mid-analysis but left
    `{ outcome: 'simple', error: ... }` SELECTED the branch and admitted it;
  - a fan-out condition that crashed with `{ error: ... }` dispatched ALL
    members, because `readActiveMembers` treats a result with no
    `activeMembers` array as "every member is active".

Renamed to `succeededResultForNode` and filtered on the committed status, so
every call site reads as what it means. Only a succeeded completion is a
decision; anything else leaves the evaluator unresolved and re-admitted.

Pre-existing kernel behaviour, but this delta touched the seam twice without
closing it — `12cc6131` scoped the facade validators to succeeded completions,
which formalized the door this leaves open on the kernel side. Both reviewer
probes are kept verbatim as regression tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(ecp-full-feature): keep the r2 reviewer evaluator-status probes as regression tests

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ecp-full-feature): reject plans whose rejoin depends on both choice branches

R2-M1. The round-2 carve-out ("a convergence node reachable from BOTH branches
is not excluded; it stays gated by its own requires") was accurate to the code
and concealed a silent deadlock. `requires` is AND semantics and exactly one
branch is ever selected, so a rejoin node downstream of both can never have all
its requires satisfied. Not excluded means never admitted; still counted by the
implicit-finish check means the Run stalls at `waiting` forever, with no
escalate and no diagnostic. My own round-2 test asserted `finish: false` as the
EXPECTED outcome, which is how the gap survived review.

Rejected at `createRuntimePlan` rather than escalated at runtime. This is a
purely static property of the plan — no Record is involved and the verdict
cannot change as a Run progresses — so the build-time check fails
`pipeline start` outright with the offending node and choice named, instead of
creating a Run that stalls and only later escalates. Runtime detection would
re-derive the same static fact on every reconcile and surface it after the fact.
Same reasoning as rejecting unresolvable fan-out/join member paths in
`3f4b0e53`: loud at build time beats silent at runtime.

Detection is transitive (a rejoin two hops below each branch is caught) and
scoped per choice, so branches of DIFFERENT choices may still converge — that
shape is satisfiable when both choices pick the converging outcomes.

Replaces the round-2 convergence test with three: direct rejoin rejected,
transitive rejoin rejected, and non-converging branches still accepted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(ecp-full-feature): spec the evaluator status rule, retry cap, and rejoin rejection

Three corrections to the delta spec, matching this round's kernel changes.

R2-M2: the rewritten Choice requirement was status-silent — it said a committed
result naming a declared outcome counts as a selection, without requiring the
completion to have SUCCEEDED. A failed evaluator can carry partial output
naming an outcome. Both the Choice and FanOut requirements now say only a
succeeded completion is a decision, with scenarios for each.

R2-m1: the unresolved-evaluator re-admission had no stated bound. Both
requirements now say re-admission is bounded and the reconciler escalates with
a code naming the evaluator, with a scenario each.

R2-M1: the requirement claimed un-selected branches "SHALL NOT block the
implicit finish", which was transitively violated by a rejoin depending on both
branches. It now states that such a plan is rejected at plan creation, and the
"convergence node is not excluded" scenario is replaced by a rejection scenario
plus one confirming that non-converging branches are still accepted.

Follow-up recorded for the ticket, deliberately NOT fixed here (out of ECP-4
scope per the round-2 review): `supported_v2_parallel` is production-unreachable
until `resolveCapabilityBindings` emits `root:`-path bindings for parallel-only
v1 shapes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-full-feature): consume parallel/1 + choice/1 in the Operations plane

proposal.md promises the parallel frontier is "visible from CLI, Management
API, and Operations". The first two planes read the projector's `parallel/1`
and `choice/1` sections; `packages/ui` had no consumer of either -- its only
section extractors were `getRootDagSection` and `getReviewCycleSection`.

Mirror the core wire types (`ParallelSectionView`, `ParallelMemberView`,
`ParallelMemberStatus`, `ParallelJoinState`, `ChoiceSectionView`,
`ChoiceBranchView` from `src/core/change-run/internal/projector.ts`) into the
UI types module, add `getParallelSection`/`getChoiceSection` alongside the
existing extractors, and render both sections in the Run detail body:
fan-out members with status/role/condition, the Join barrier state, budget
usage, the projected member frontier, and the choice outcome with each
branch marked active or inactive.

`joinPath` and `outcome` are mirrored as OPTIONAL properties: the server type
is `string | undefined` and JSON drops undefined keys on the wire.

Every value is read from the projection. The UI recomputes no member status,
join state, count, or branch activation -- the same discipline it already
follows for the root-DAG frontier.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(ecp-full-feature): Operations-plane parity for parallel/choice (8.6)

Task 8.6 asks for parity across CLI, Management API and Operations. The first
two were covered by `test/core/change-run/projector-parallel-choice.test.ts`;
the task was ticked with an annotation admitting the UI plane was missing.

Add the fourth plane under `packages/ui/test/`, per the existing
`cross-plane-parity.test.tsx` convention: render the REAL `OperationsSection`
against the SAME fixture the node-side parity test builds (FanOut
`root:experts` cap 2 / budget 3 with the condition committed and the required
member settled; Choice `root:pick` committed `simple`) and assert the DOM
carries the projected member statuses, roles, conditions, join state, budget
usage, frontier counts, key blockers, outcome and branch activation.

The canonical section constants are the verbatim `projectRunView` output for
those Records, not values written to match the component; their provenance is
documented in the file header.

Three assertions feed deliberately INCOHERENT sections -- counts that no tally
of `members` produces, a `waiting` join over all-succeeded members, and an
active branch that is not the selected outcome. With coherent fixtures a UI
that recomputed these is indistinguishable from one that consumes them; the
disagreement is what pins the behaviour. Verified by mutation: recomputing
`activeCount` from member statuses and branch activation from `outcome` turns
those tests red.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ecp-full-feature): projector must read the LAST SUCCEEDED evaluator result

R3-M1. The `succeededResultForNode` sweep closed the kernel but missed the
projector, which had a second defect on top: it took the FIRST result-bearing
attempt where the kernel takes the LAST, and ignored completion status
entirely.

Both defects only matter together, and the retry cap is what made them live.
After a failed-then-retried-successfully evaluator — now an ORDINARY path
precisely because retries exist — the projection kept reading the failed first
attempt forever: members `waiting`, joinState `not-reached`, while the Run was
actually executing. It also displayed a crashed evaluator's partial output as a
real selection. Same shape as N1: a fix makes a dead path live, and the newly
live path carries its own bug.

`succeededEvaluatorResult` mirrors the kernel reader — last attempt, succeeded
only, object-shaped — and both section builders use it. Display-only; execution
semantics are untouched.

All three parity planes share this one reader, which is exactly why the parity
suite could not catch it: every plane was wrong identically.

Reviewer probes in `reviewer-r3-ui-constants-provenance.test.ts` kept verbatim.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(ecp-full-feature): evaluator retry cap must not count the in-flight attempt

R3-m1. `occurrenceForNodeId` counts EVERY invocation including one still
active, so with 2 failures committed and the third attempt in flight, the next
reconcile — any settle, a `resume-run`, or another node completing — escalated
`choice_evaluator_unresolved` and terminated a Run that was about to resolve,
discarding live paid work.

`evaluatorAttempts` reports both facts separately. While an attempt is active
no verdict is due at all: the pass emits neither an admit (which would
double-dispatch) nor an escalate. Only committed attempts count against the cap.

The round-3 cap tests passed honestly because their records are all-terminal;
they still do.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(ecp-full-feature): keep the r3 reviewer probes as regression tests

3 probes that proved R3-M1/R3-m1 plus 4 provenance tripwires pinning the UI
parity constants to the real projection.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat(ecp-product-closure): Operations consumes review-cycle/1 + kernel-anchored parity constants

ECP-1's shipped `executable-review-cycle` delta requires CLI, Management API
and Operations to consume the same ChangeRunView review-cycle section. Two of
the three planes did; `getReviewCycleSection` had zero consumers in
`packages/ui/src`. Adds a `ReviewCycleSection` renderer to OperationsSection
following ECP-4's parallel/choice pattern: round against cap, phase, outcome,
findings with severity + status, bound actors, wait reason - every value read
from the projection, none re-derived (the round cap, the cle…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant