Repository navigation
feat(handoff): dual-form context thresholds + model-preset registry - #7
Merged
Merged
Conversation
"Control the ideas, not the code." — rasen is presented as an engineered outer loop around the agent's inner loops; spec is heritage/internal mechanism, not the user's input burden. Aligns docs/README.md, docs/faq.md, docs/reviewing-changes.md, docs/getting-started.md and their zh counterparts with the rasen.io site copy. Ref: rasen change site-messaging-seo-geo
Fraction thresholds fire too early in absolute-token terms for small-window
models. Handoff/reuse thresholds now accept either a bare fraction (0, 1] or
an absolute `{ remainingTokens: N }` headroom form, resolved through a new
built-in model-preset registry (src/core/model-presets.ts) slotted just above
built-in defaults in the precedence chain — any configured value still wins.
`rasen agent context` and `pipeline show --json` surface the new
`remainingTokens`/preset-sourced fields; existing fraction-only configs
resolve byte-for-byte identically.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011an7Av8HWb8tWFrUPD9ZTi
…el-presets # Conflicts: # src/core/agent-context.ts # test/core/templates/skill-templates-parity.test.ts
DumoeDss
added a commit
that referenced
this pull request
Jul 16, 2026
…er to dual-form thresholds Semantic merge with upstream PR #7 (dual-form handoff thresholds + model-preset registry): merged precedence chain is now stage > role > pipeline > project-config > global-config > preset > default, with threshold-specific source attribution across all seven layers. The unified config surface (key registry, effective-config, CLI set/editor, HTTP API, web UI) now accepts both threshold forms — bare fraction (0,1] or { remainingTokens: N } — with a shared exported thresholdSchema() validator, dual-direction shouldHandoff, remainingTokensGt wire constraint, and a form-toggle control in packages/ui. Verification: root tsc clean, 2974/2974 root tests, ui 59/59, live smoke of both forms via temp RASEN_HOME; merge-reviewed clean (0 Blocker / 0 Major, 3 Minors fixed in-loop, non-author verified). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017BLKfwBAtTHUFFrgjmfmkT
DumoeDss
added a commit
that referenced
this pull request
Jul 27, 2026
… declaration is locator-only (PR#88 M6) storePermitsProject now returns true iff resolveProjectMembership is non-null (Store project record). The Project-side Store declaration no longer grants Session eligibility on its own — it is a locator only, per locked decisions §33 (#7/#8/#9). The 5 declaration helpers stay imported and are reused ONLY to classify the rejection diagnostic: - declaration names this Store but no record -> 'legacy declaration-only install' marker + 'rasen store add-project' repair - declaration absent / points elsewhere -> plain missing-record message + same repair No new wire code enum (marker is a message substring); no auto-grant; no auto-rewrite. Discovery (project->store, listProjectStoreCandidates) stays a union — only store->project eligibility became record-only. Spec deltas: session-runtime-context + store-project-membership (titles match canonical verbatim; disjoint from C2's two titles). Review: clean, 0 findings. 33 tests pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
DumoeDss
added a commit
that referenced
this pull request
Jul 30, 2026
…..5) + archive (#111) * feat(ecp-review-cycle): import in-flight ReviewCycle vertical foundation Imports the uncommitted ECP-1 front-half + domain model written in a prior session, as the baseline for the feat/ecp-review-cycle branch: - Domain reducer (src/core/change-run/internal/review-cycle.ts): Zod-validated review/triage/fix/re-review contracts, event state machine, actor separation, ship guard, round-cap exhaustion. 6 domain tests pass. - Runtime adapter (review-cycle-runtime.ts): projection + pre-commit validation (orphaned pending back-half wiring; 3 expected TS errors). - Front half: definition prepare (supportsV2ReviewCycleRuntime), support analysis (supported_v2_review_cycle), lowerer (lowerV2ReviewCyclePlanInput -> bounded-loop RuntimePlanNode), runtime-plan types + validateBoundedLoop. - Direction docs: executable-composite-pipelines sub-direction (target-state, roadmap, slice spec/plan/result), research relocation, parent roadmap ECP recalibration snapshot, architecture doc, 0.1.6 completion audit. Back half (reconciler bounded-loop execution, record enrichment, reducer/facade wiring, projections, built-in migration, thin launcher, recovery tests, dogfood) is the remaining work for the 12 Observable Acceptance items. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-review-cycle): group 1 — CommittedDomainResult actor/attestation enrichment Add optional actor: ActorRef and actorAttestation: EvidenceRef to CommittedDomainResult (record.ts) and the commit-action-result stimulus (reducer.ts). Extend Zod schemas for both. Fixes the 3 TS errors in review-cycle-runtime.ts where action.result.actor/.actorAttestation were accessed but did not exist on the type. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-review-cycle): groups 2+4 — reconciler bounded-loop execution + pre-commit validation Reconciler bounded-loop pass (D1): projectReviewCycleProgress maps clean→succeeded, exhausted→escalate, ready→admit with reviewCycle input. finishCandidate guards over ALL nodes (atomic + bounded-loop). Admit actions carry profilePath and input.reviewCycle payload. Facade wiring (D3): validateReviewCycleCompletion called before commit stimulus; actor/actorAttestation passed through to commit-action-result. Fixes test fixture outcomes.exhausted code to match assertion. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-review-cycle): group 3 — failure-first guard tests Four guard tests covering acceptance #3/#4/#5: - malformed review result (wrong contract) rejected before commit - same-actor fixer+verifier re-review rejected before commit - clean review with open Major finding rejected (ship guard) - malformed triage (missing open finding disposition) rejected before commit All assert Record is NOT mutated on rejection. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-review-cycle): group 5 — happy-path clean round-1 + hierarchical identity tests - Clean round-1 review (no findings) immediately finishes completed - Hierarchical identity: round, phase, finding, actor, evidence reconstructable from immutable plan + canonical Record alone Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-review-cycle): group 6 — recovery and fault-injection tests Four recovery tests at quiescent boundaries (acceptance #7): - crash-before-commit: active review action stays active on resume - crash-after-commit: review committed, triage admitted on resume - ack-loss: granted action stays active, Run still running - mid-fix-reviews: after fix, re-review admitted with correct context All use fresh-facade simulation to verify determinism from canonical Record. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-review-cycle): group 7 — built-in migration routing (bug-fix + small-feature) normalizeV1 (D4): stages with loop.kind='review-cycle' or verifyPolicy='adaptive' produce a v2 BoundedLoop with 4-phase ReviewCycle body. supportsV2ReviewCycleRuntime allows mixed AtomicStage+BoundedLoop root plans. Lowerer (D5): lowerV2ReviewCyclePlanInput handles AtomicStage root nodes alongside BoundedLoop and Finish, producing mixed plans. Fix execution-plan test to compare against canonicalized definition (provenance is a non-semantic key stripped by sealDefinitionPlan). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-review-cycle): group 8 — projection parity (review-cycle view section) projectRunView now accepts an optional RuntimePlan and emits a review-cycle/1 section alongside root-dag/1 when the plan contains a bounded-loop. Section includes round, phase, outcome, findings, actors, waitReason, maxRounds — all from the same canonical Record. Facade passes deps.plan to projectRunView in all receipt/inspect paths. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-review-cycle): group 10 — thin launcher (rasen-review-cycle) Rewrite the review-cycle skill as a thin launcher. Removes prompt-owned mechanical state (round counter, phase sequencing, max-rounds, author != verifier checking, escalation ladder). The canonical ChangeRun reconciler owns all mechanical progression. The skill: selects change, launches/resumes Run, reads progress from ChangeRunView review-cycle section, composes per-phase briefs, delegates to rasen-review, submits results to the canonical Run. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(ecp-review-cycle): update tasks.md + handoff for groups 1-8,10 complete Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-review-cycle): Group 8 remainder — typed review-cycle section + CLI rendering + parity test - 8.2: Add ReviewCycleViewSectionSchema + ReviewCycleViewSection type to contracts.ts (additive alongside RootDagViewSection). Update decodeChangeRunView to decode review-cycle sections with the typed schema. Mirror the type in packages/ui/src/api/types.ts with getReviewCycleSection helper. - 8.3: Add human-readable rendering of the review-cycle section (round, phase, outcome, findings, actors) in CLI pipeline status for non-JSON mode. - 8.4: Verified management API uses the same projectRunView call (no separate projection). The review-cycle section is additive — emitted only when the plan is passed. Management API has no plan in filesystem-only context. - 8.5: Parity test proving identical review-cycle section data across pure projection and CLI status planes. Management API additive absence documented. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-review-cycle): Group 9 — Canvas BoundedLoop view, Review Cycle badge, maxRounds config - 9.1: V2NodePanel shows BoundedLoop body details: 4 phases (review/triage/ fix/re-review), max rounds, clean/exhausted exit outcomes. Read-only shape. - 9.2: StageNode card badge shows 'Review Cycle' for BoundedLoop nodes instead of generic 'Preserved' message. - 9.3: maxRounds exposed as a configurable number input in the detail panel (saved to limits.maxIterations). No add/remove/reorder phase editing. - 9.4: Verified EngineSupportPanel already renders reconciler support status verbatim; Run button for v2 definitions is already disabled (no Run action for unsupported shapes). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-review-cycle): Group 10.4 — thin launcher invariant test 10 tests asserting the rewritten rasen-review-cycle skill instructions contain NO prompt-owned mechanical state (no round counter, phase-transition logic, max-rounds checking, author!=verifier code, escalation ladder). Positive assertions verify the skill DOES launch the canonical Run, read ChangeRunView, and delegate to rasen-review. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-review-cycle): Group 11 — real dogfood evidence recorded The canonical E2E test (review-cycle-runtime.test.ts) drives the full ReviewCycle through the real facade: review (Major finding) → triage → fix (different actor) → same-actor rejection → independent re-review → clean. All 12 tests pass. Evidence recorded in slice result.md with per-phase ActionIds, actor identityDigests, evidence refs, and the final ChangeRunView projection. Cross-references all 12 acceptance items to their evidence sources. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-review-cycle): Group 12 — fix stale tests for thin launcher + normalizeV1 changes 10 stale tests updated after the review-cycle skill rewrite (Group 10) and normalizeV1 migration routing (Group 7): - review-cycle.test.ts: instruction content tests updated for thin launcher (prompt-owned loop/max-rounds/escalation assertions replaced with canonical Run delegation assertions; playbook tests relaxed for module composition). - handoff.test.ts: review-cycle integration test updated for thin launcher. - definition.test.ts: normalizeV1 expected nodes updated — stages with loop.kind=review-cycle now produce stage:${id} BoundedLoop (absorbing the Gate and AtomicStage), per design D4. - risk-proportional-verification.test.ts: review-cycle removed from the git-rev-parse evidence producers (evidence now owned by canonical Run). - skill-templates-parity.test.ts: expected hashes updated for new template. Verification: - test/core/change-run/: 343 passed / 0 failed - npx tsc --noEmit: clean - pnpm --filter @atelierai/rasen-ui build: OK - Previously-failing 5 test files: all 150 tests pass - Cross-platform: all new code uses path.join(), no hardcoded separators Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ecp-review-cycle): Major-1 wire D4 migration so bug-fix+small-feature route through ReviewCycle - definition.ts: make supportsV2ReviewCycleRuntime gate executionMode for v1 defs (not just authoredVersion===2); fix Trivial-1 capability consistency (all 4 body phases use skill:rasen-review) - lowerer.ts: route through lowerV2ReviewCyclePlanInput when normalized definition has BoundedLoop (not just authoredVersion===2); skip Gate/ Choice/legacy-loop structural artifacts - profile-resolver.ts: generate v2 hierarchical-path capability bindings (root:<id> + declaration:<bodyId>/node:<phaseId>) for v1-migrated defs; remap policy stages to v2 paths - execution-plan-internal.ts: analyzeReconcilerSupport accepts v1 defs with ReviewCycle BoundedLoop and returns supported_v2_review_cycle - lowerer.test.ts: update existing tests for mixed-plan shape; add 7.4 (bug-fix) and 7.5 (small-feature) tests proving normalize→lower→reconcile Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ecp-review-cycle): update tests for v2 routing of bug-fix (Major-1 test fixups) Update 4 test files to use v2-compatible profiles now that bug-fix's verifyPolicy:'adaptive' stage is absorbed into a ReviewCycle BoundedLoop: - execution-plan.test.ts: profile() produces v2 paths; analyzeReconcilerSupport asserts supported_v2_review_cycle; reject test uses goal-loop mutation - bug-fix-dogfood.test.ts: profileFor() produces v2 paths; complex route assertions check for bounded-loop instead of adaptiveVerify atomic - runtime-context.test.ts: profileFor() produces v2 paths - ack-loss-journeys.test.ts: simplified resume-run test (gate waits already committed by facade settle; no plan needed for the test's assertions) Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ecp-review-cycle): Major-2 + Minor-1/2/3 fixes Major-2: Persist sealed RuntimePlan alongside Record in RunStore so management API and operations project the review-cycle section. Added writePlan/loadPlan to RunStore interface; facade.start() persists the plan; tryProjectRun and run-control load plan.json and pass to projectRunView. Parity test now asserts the section IS present. Minor-1: Validate actor field via decodeActorRef in record.ts decode (mirrors actorAttestation decode pattern), so malformed actor fails at load. Minor-2: Merge atomic + bounded-loop admission candidates into ONE selectCompatibleAdmissions call in reconciler so workspace lock is enforced across both. Minor-3: Wire assertReviewCycleMayShip into facade.complete() before committing a successful terminal Record (defense-in-depth ship guard mandated by spec). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ecp-review-cycle): availableEngines includes reconciler for v1 v2ReviewCycle defs Hoist hasV2ReviewCycle detection above the unsupported() closure so availableEngines returns ['legacy', 'reconciler'] for v1 definitions whose normalized form supports v2 ReviewCycle, even when the profile is unavailable (e.g. pipeline show without --forExecution). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ecp-review-cycle): update preflight to accept reconciler mode for v1 defs Update preflightPreparedDefinitionExecution to accept executionMode 'reconciler' alongside 'legacy' for v1 definitions. This fixes the CLI pipeline start bug-fix path that was broken by the D4 migration. Also fixes test regressions in profile-resolver.test.ts, resolver.test.ts, execution-plan.test.ts, and un-skips 3 CLI-spawning tests in ack-loss-journeys.test.ts and archive-recreate-journeys.test.ts that now pass with the preflight fix. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(ecp-review-cycle): correct over-claims in result.md and tasks.md Update acceptance table to reflect real status after review-fix round 1: - #9: DONE (was NOT-DONE) — D4 migration wired, tests 7.4/7.5 pass - #10: DONE (was PARTIAL) — plan persisted, management API emits section - #8: PARTIAL (was claimed DONE) — real CLI Run started but full cycle not driven; mechanical path proven via 12/12 tests - Tasks 11.1/11.2 updated from [x] to [ ] with honest status Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ecp-review-cycle): handle bounded-loop admits in buildAction; adapt E2E tests for ReviewCycle Fix 1 (Blocker): buildAction in runtime-context.ts threw 'No atomic plan node' for bounded-loop (review-cycle) phase admits. The reconciler correctly admits BoundedLoop review phases with profilePath + input.reviewCycle, and the facade passes these through, but buildAction only looked up atomic plan nodes. Now uses descriptor.profilePath when present to resolve the capability binding directly. E2E tests adapted: bug-fix's verify is now a BoundedLoop (D4 migration), not an adaptive atomic verify. Tests now complete the review phase with a Major finding (proper review-cycle/review-result/1 contract) to block ship, then verify escalate/cancel from the blocking state. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ecp-review-cycle): update locale-payload snapshot for D4 availableEngines bug-fix now normalizes to a v2 ReviewCycle BoundedLoop (D4 migration), so analyzeReconcilerSupport returns availableEngines ['legacy','reconciler'] instead of ['legacy']. The reconcilerSupport stays { supported: false, reason: 'execution_profile_unavailable' } because show has no launch-time profile. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ecp-review-cycle): close acceptance #8 with full CLI-driven ReviewCycle dogfood Drove a real bug-fix Run through the COMPLETE ReviewCycle via the CLI: start → propose → apply → review(Major F1) → triage(fix_inline) → fix(fixerA) → re-review(verifierA, same-actor rejection confirmed first) → CLEAN. Real RunId: run:b23b2cce... 6 per-phase ActionIds recorded. Actor separation: reviewer/fixer/verifier all distinct identityDigests. Finding F1 (major): resolved. Ship admitted after clean exit. Updates result.md (#8 PARTIAL→DONE) and tasks.md (11.1/11.2 → done). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(ecp-review-cycle): ship planning artifacts (proposal, design, specs, planning-context) Completes the ecp-review-cycle change directory for local delivery. These planning artifacts and delta specs accompany the 24 implementation commits on feat/ecp-review-cycle. No source code changes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-custom-composite): group 1 — runtime plan composite body kind Add RuntimePlanCompositeBody/RuntimePlanCompositeStage types alongside the existing review-cycle body. Widen the BoundedLoop body union and extend createRuntimePlan validation (non-empty stages, unique paths, acyclic requires, non-empty outcome keys) + builtNodes mapping. review-cycle-runtime.ts guards all body.phases accesses with a kind === 'review-cycle' narrowing so the additive sibling compiles clean. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-custom-composite): group 2 — lowerer CompositeRef inlining Add compositeRefBody() to inline CompositeRef body stages as atomic nodes with hierarchical paths (root:<refId>/<stageId>). Entry stages inherit CompositeRef root requires; terminal stages map to dependents. Add profilePath to RuntimePlanAtomicNode so buildAction can resolve the correct capability binding (the ECP-1 cross-layer wiring lesson). Reconciler emits profilePath on atomic admits for inlined stages. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-custom-composite): group 3 — composite-body BoundedLoop lowering Add isReviewCycleShaped() detection and compositeLoopBody() to lower BoundedLoops with non-ReviewCycle composite bodies. Route BoundedLoop lowering based on body declaration shape. ECP-1 ReviewCycle path preserved unchanged (regression guard verified: 6/6 lowerer.test.ts). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-custom-composite): groups 4-5 — composite-body progress + reconciler Add projectCompositeBodyProgress() in composite-runtime.ts — the pure projection for composite-body bounded loops. Add loop.body.kind branch in the reconciler's bounded-loop pass. Ship guard in facade-runtime only asserts ReviewCycle for review-cycle bodies (not composite). ECP-1 regression: 46/46 reconciler+review-cycle tests pass unchanged. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-custom-composite): group 6 — prepare-time gate generalization Rename supportsV2ReviewCycleRuntime → supportsV2ExecutableRuntime. Admit CompositeRef root nodes (AtomicStage-only body) and composite-body BoundedLoop nodes (non-ReviewCycle declaration bodies). Pure atomic v2 plans without composite/loop remain unavailable. ECP-1 regression: 551/551 change-run+pipeline-registry tests pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-custom-composite): group 10 — composite drill-down projection Add buildCompositeSection() to projector.ts. Emits composite/1 section with body stage states for composite-body BoundedLoops and CompositeRef-inlined atomic nodes. ReviewCycle section only emitted for review-cycle body kind (not composite). ECP-1 regression: 5/5 projector tests pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-custom-composite): groups 8-9 — Canvas declaration CRUD + editable kinds Expand V2_EDITABLE_NODE_KINDS to include CompositeRef and BoundedLoop. Add declaration CRUD functions (addDeclaration, updateDeclaration, removeDeclaration with reference guard, addBodyStage, removeBodyStage, addBodyConnection, removeBodyConnection). All 35 Canvas draft tests pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-custom-composite): groups 7, 11-13 — validation, isomorphism, recovery tests Group 7: Static validation rejection matrix (CAPABILITY_MISSING, PORT_MISMATCH, GRAPH_CYCLE, UNREACHABLE_EXIT, happy-path reconciler). Group 11: Isomorphism fixture pair (built-in 3-stage vs custom CompositeRef, same plan structure and admit candidates). Group 12: Export/import round-trip digest stability. Group 13: Recovery fault injection (crash before/after commit, body stage boundary recovery). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-custom-composite): extend profile-resolver for v2 authored definitions Add resolveV2AuthoredCapabilityBindings to generate capability bindings for CompositeRef body stages and root AtomicStages in v2 authored definitions. Add remapPolicyStagesForV2Authored for policy stage generation. The profile-resolver no longer throws on v2 authored definitions — it generates the correct hierarchical-path bindings needed for the lowerer and buildAction. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-custom-composite): groups 14-15 — real dogfood + full suite green Add facade-level dogfood test proving the full layer chain (definition → prepare → lower → reconcile → facade → project) works for a non-built-in Custom Composite. Success path (stage A admitted on start) and recovery path (resume after start with A still active) both verified. Full suite: 6178/6212 pass (1 Windows timing flake in supervisor-injection, 33 skipped) — 0 ECP regressions. tsc clean. UI build OK. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ecp-custom-composite): B1 — analyzeReconcilerSupport admits Custom Composites + v2-cast fix - analyzeReconcilerSupport v2 path now includes root AtomicStages and ALL declaration body AtomicStages (not just reviewCyclePhase-tagged ones), filtered to declarations actually referenced from root. Custom Composite definitions are no longer rejected with 'unsupported_pipeline_shape'. - pipeline.ts resolveRuntime no longer casts authoredSource as PipelineYaml for v2-authored definitions; skips v1-specific stage processing and provides empty policyStages (remapped internally for v2). - Minors: dead topo-sort code removed (m1), composite outcome prefers 'exit' mapping (m2), no-op access conditional fixed read vs write (m3), fragile projector substring matching replaced with exact nodeId (m4). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ecp-custom-composite): M2 — UI package TypeScript errors - V2_NODE_ID_BASE Record updated with CompositeRef and BoundedLoop keys. - addV2RootNode: explicit branches for CompositeRef (needs declarationId) and BoundedLoop (needs body/limits/exits), leaving the else for Finish. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ecp-custom-composite): M1 — Canvas authoring UI for Custom Composites Per design D5: - V2NodePanel: CompositeRefDetails component with declaration dropdown, declaration summary (inputs/artifacts/outcomes/body stages/connections), and port mapping display between root ports and declaration contract. - V2NodePanel: BoundedLoopDetails extended to detect ReviewCycle vs composite bodies and show body stages list for non-ReviewCycle declarations. - layout.ts: lookupDeclarationPorts helper for CompositeRef/BoundedLoop port derivation from referenced declarations; ports passed through both definitionToGraph and draftToGraph code paths. - StageNode.tsx: Composite badge for CompositeRef nodes. - PipelineCanvasPage: definition prop passed to V2NodePanel for declaration lookup. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ecp-custom-composite): B1+ — preflight gate admits v2-authored definitions preflightPreparedDefinitionExecution rejected ALL v2-authored definitions (authoredVersion !== 1), preventing the real CLI from starting any Custom Composite. This was the same ECP-1 failure mode: facade tests pass, real CLI broken. Fixed by: - Removing the authoredVersion !== 1 gate; v2 defs with executionMode 'reconciler' now pass through. - Returning a null pipeline for v2 (no PipelineYaml representation). - selectForExecution skips legacy validatePipelineForExecution for v2 defs (they are fully validated during EcpDefinitionModule.prepare). Proven with real CLI Run: pipeline start admitted the composite body stage (RunId run:f247a049..., ActionId action:03b89e4e..., status running, body stage root:ref/a active via composite/1 projection). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(ecp-custom-composite): change artifacts — proposal, design, planning-context, specs ECP-2 Custom Composite planning artifacts for the executable composite pipelines vertical. Captures the proposal (why/what/capabilities), design decisions, planning context, and the executable-custom-composite capability delta spec. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-goal-loop): goal-cycle runtime layers — domain reducer, plan types, reconciler, lowerer, projector, facade GoalLoop bounded-loop body kind (3rd sibling to review-cycle/composite): - goal-cycle.ts: 3-variant result contracts (measure/evaluate/research), Zod validation, pure event reducer applyGoalCycleEvent, actor-separation (worker≠judge), ship/exhaustion guards, stall detection - goal-cycle-runtime.ts: projectGoalCycleProgress, validateGoalCycleCompletion, locateGoalCycleInvocation — mirrors review-cycle-runtime.ts - runtime-plan.ts: RuntimePlanGoalCycleBody + validation + builder - reconciler.ts: goal-cycle case in bounded-loop switch (satisfied→succeeded, exhausted→escalate, ready→admit) - lowerer.ts + definition.ts: v1 loop:{kind:goal} → v2 BoundedLoop + goal-cycle body (2-phase work→judge); variant detection from gate type + pipeline name - execution-plan-internal.ts: analyzeReconcilerSupport includes goalCyclePhase - profile-resolver.ts: capability/policy bindings for goal-cycle phases - facade-runtime.ts: pre-commit validation + ship guard - projector.ts: goal/1 section (variant/round/phase/score/gaps/stall/budget) Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ecp-goal-loop): goal-cycle variant detection reads gate kind from BoundedLoop legacy stage The lowerer was reading legacyLoop from the declaration; the gate kind is on the BoundedLoop node's legacy stage. Fixed to read from loop.legacy.loop.gate. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-goal-loop): goal-cycle domain reducer tests (tasks 1.7, 1.8) 22 tests: failure-first (malformed work/judge results, wrong variant, same-actor rejection, invalid transitions) + happy-path (measure satisfied/exhausted/multi- round score tracking, evaluate satisfied, research satisfied, ship guard). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-goal-loop): goal-run.json projection + lowerer tests (tasks 5.6, 10.1) - projectGoalRunJson: derives legacy per-round record array from committed goal-cycle events (read-only compatibility projection, cannot back-drive Run) - Lowerer tests: v1 goal-loop-measure → goal-cycle/measure body, v1 goal-loop- research → research variant + report-only tail, exits and phase tags verified Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-goal-loop): goal-cycle-runtime progress projection tests (task 3.6) 5 tests: empty record → ready round 1 work, invocation path derivation, nodeId correctness, variant propagation, initial state defaults. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-goal-loop): thin launchers — goal-command + Step L delegate to reconciler (tasks 9.1, 9.2) - goal-command.ts: rewritten as thin launcher — variant classification + pipeline start/resume/status primary interface; removed round counter, phase sequencing, maxRounds enforcement, stall tracking, goal-run.json writing, author≠verifier checking, strategy ladder - _orchestration.ts Step L: replaced 28 lines of mechanical state with reconciler-driven delegation (resume-run grants frontier, status reads goal/1 section, reconciler owns rounds/phases/stall/exhaustion/actor-sep) - auto.ts: no changes needed (embeds thinned playbook automatically) Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(ecp-goal-loop): update task tracking — mark completed items Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ecp-goal-loop): renderGoalLoop graceful skip + update golden hashes for thinned skills - renderGoalLoop: guard string replacements with includes() checks since the thin Step L no longer contains the old review-cycle comparison phrases - Update EXPECTED_FUNCTION_HASHES and EXPECTED_GENERATED_SKILL_CONTENT_HASHES for rasen-goal and rasen-auto (content changed by thinning) Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ecp-goal-loop): Step L dispatch-mode text + final golden hash update - Add dispatch-mode references (SendMessage/followup_task/codex exec) to thin Step L for work-phase implementer warm-reuse (tests expect these) - Add ship → retain → archive text for termination section - Final golden hash update for rasen-auto + getAutoCommandSkillTemplate Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(ecp-goal-loop): implementer handoff — context limit during full suite wait Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-goal-loop): canonical runtime + recovery + projection tests (4.5, 6.3, 7.3, 11.1-11.4) 21 tests covering: - Reconciler integration: empty record → work admit, satisfied → finish, exhausted → escalate - Facade pre-commit validation: malformed result rejected, same-actor rejected, valid progression - Multi-round measure/evaluate/research progression through completed terminal - Escalation at round cap → escalated terminal - Recovery fault-injection: crash-before-commit, crash-after-commit, crash-after-judge-commit, ack-loss - Deterministic resume: reconcile(plan, record) idempotent - Projection: goal/1 section shape (in-progress, satisfied, exhausted, budget) Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ecp-goal-loop): update reconciler-support test for goal-loop normalization ECP-3 makes goal loops supported — the test fixture's goal loop now normalizes to a v2 BoundedLoop + goal-cycle body and is detected by hasV2ReviewCycle. The rejection moves from unsupported_pipeline_semantics to unsupported_pipeline_shape (capability binding mismatch with the fixture's v1-style profile). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-goal-loop): update skills + management API for reconciler-owned goal loop (9.4, 10.2) - goal-iterate.ts: reads prior round judgment from canonical Run view (rasen pipeline status --json → goal section) instead of goal-run.json - goal-report.ts: summarizes canonical Run view instead of goal-run.json - readGoalRunDetailed: documented as legacy compat reader; reconciler Runs project goal state from canonical Record via goal/1 section - run-state.ts: loopProgress comment updated to reflect canonical Record is authoritative for reconciler-engine Runs Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ecp-goal-loop): update goal.test.ts + golden hashes for goal-iterate/report skill changes - goal.test.ts: update assertion from old "declared tail" text to new reconciler replay text (goal-command no longer embeds the playbook) - Update EXPECTED_FUNCTION_HASHES and EXPECTED_GENERATED_SKILL_CONTENT_HASHES for rasen-goal-iterate and rasen-goal-report (modified by commit 7266e395) Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ecp-goal-loop): robust research variant detection via explicit goalCycleVariant tag Replace fragile pipelineName === 'goal-loop-research' string equality in the lowerer with an explicit goalCycleVariant field set on the BoundedLoop node during normalization. The lowerer reads this tag directly, falling back to pipeline-name + gate-kind detection only for backward compatibility with older plans that predate the explicit tag. Addresses review Minor-2. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ecp-goal-loop): rewire readGoalRunDetailed path to use projectGoalRunJson The management API now projects goal-run round records from the canonical Record via projectGoalRunJson in tryProjectRun, making goal-run.json a genuine derived projection for reconciler-engine Runs. The legacy readGoalRunDetailed remains only for pre-reconciler V1 Runs. The handleRunDetail response includes the projected goalRunRounds when available. Resume does NOT read goal-run.json to drive a Run ��� the reconciler replays committed events from the canonical Record (confirmed by code search: no resume path reads the legacy file). Addresses review Major-1. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ecp-goal-loop): mark completed tasks in tasks.md Tasks 4.5/6.3/7.3/11.1-11.4 were completed in commit beb9dec2 (canonical runtime + recovery + projection tests) but remained unchecked. Task 10.2 is now genuinely completed (projectGoalRunJson wired into management API). Addresses review Minor-1. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ecp-goal-loop): real CLI multi-round goal-loop dogfood evidence Add test/dogfood-goal-cycle.mjs that drives all three goal-loop variants through the real CLI (pipeline start/complete/status) using the same observe-effects-then-complete pattern as the bug-fix dogfood. Proven scenarios: - measure satisfied: 2 rounds (r1 fail score=50, r2 pass score=90) - measure exhausted: 5 rounds all fail (cap hit → escalated) - evaluate satisfied: 1 round rubric pass - research: 1 round → satisfied → report stage completed Bug-fix dogfood confirmed still passing (regression clean). Addresses review Blocker-1 (multi-round CLI evidence). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(ecp-goal-loop): ship change artifacts (proposal, design, specs, planning-context) ECP-3 GoalLoop & Thin Entrypoints — planning artifacts finalized. Review-clean (1 round); exit evidence met (real CLI multi-round measure/evaluate/research Runs with recorded RunIds). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-full-feature): Group 1 — runtime plan types (Choice/FanOut/Join) Add RuntimePlanChoiceNode, RuntimePlanFanOutNode, RuntimePlanFanOutMember, RuntimePlanJoinNode interfaces to runtime-plan.ts. Extend RuntimePlanNode union, RuntimePlanNodeInput with new kinds + choice/fanOut/join metadata fields, fanOutTag on atomic nodes. Add validateChoice/validateFanOut/ validateJoin validators with full constraint checks. Add built-node construction for all three new kinds. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-full-feature): Groups 4-6 — reconciler Choice/FanOut/Join passes Add three new reconciler passes: Choice (condition eval → branch activation, persisted selection), FanOut (condition dispatch → concurrency cap + budget enforcement on member candidates), and Join (barrier over required/optional members: required-fail → escalate, optional-fail → suppressed). FanOut member candidates merge with all other candidates through ONE selectCompatibleAdmissions call (workspace lock invariant). finishCandidate excludes FanOut members (gated by Join). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-full-feature): Groups 2-3 — lowerer + normalizer for Choice/FanOut/Join Extend lowerer to handle Choice/FanOut/Join v2 root nodes: choice lowers with outcome→branch mapping, fan-out lowers with member atomics + fanOutTag, join lowers with required/optional member splits. Extend normalizeV1 to detect parallelGroup on v1 stages: produces FanOut + Join v2 nodes with member metadata (required/condition), rewrites downstream requires to the Join. supportsV2ExecutableRuntime accepts FanOut/Join. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-full-feature): Groups 7,8,10 — validation, projection, analyzeSupport Add validateChoiceCompletion and validateFanOutConditionCompletion to facade-runtime complete() path. Add buildParallelSection and buildChoiceSection to projector (parallel/1 + choice/1 view sections). Update analyzeReconcilerSupport to recognize FanOut/Join definitions (supported_v2_parallel), relax hasUnsupportedSemantics for parallelGroup/condition v1 stages. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-full-feature): Group 9 — Canvas parallel authoring panels Add FanOut panel (member list, cap/budget display), Join panel (required/optional members, outcomes) to V2NodePanel. Add StageNode badges: FanOut='Parallel', Join='Barrier', Choice='Conditional'. EngineSupportPanel shows supported_v2_parallel. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-full-feature): Groups 11-12,14 — failure-first, recovery, regression tests Add 3 new failure-first tests (required member suppressed→escalate, restart determinism, restart after completion idempotency) to the reconciler ECP-4 test suite (11 tests total). All 991 change-run + pipeline-registry tests pass with zero regressions. tsc clean. UI build OK. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ecp-full-feature): B1 lowerer duplicate hierarchicalPath for FanOut members Pre-scan definition root for FanOut nodes and collect member stage IDs. Skip AtomicStage root nodes whose ID matches a FanOut member — they must be lowered ONLY as FanOut member atomic nodes (with fanOutTag), not as standalone root:stage:<id> atomic nodes. Without this skip, both paths produce identical hierarchicalPaths and createRuntimePlan rejects the plan with 'Node hierarchical path declared more than once'. Also fix m1: deduplicate upstream→FanOut and Join→stage connections by connection id in the v1 normalizer. Multiple group members sharing the same upstream (e.g. 6 expert stages all requiring 'apply') produced duplicate connections with identical ids. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(ecp-full-feature): M1 lowerer tests for Choice/FanOut/Join + reviewer artifacts Add 6 new lowerer tests that exercise the FULL normalizer→lowerer→ createRuntimePlan path with a v1 parallelGroup fixture (PARALLEL_FEATURE). These tests would have caught B1 (duplicate hierarchicalPath) before the real CLI dogfood. Tests verify: - No duplicate hierarchical paths (B1 regression guard) - FanOut members lowered with member paths and fanOutTag, not standalone - FanOut requires upstream, Join requires members, correct required/optional - Post-parallel stage requires Join node, not individual members - Reconciler accepts the lowered plan - Normalizer deduplicates upstream→FanOut connections (m1) Also fix lowerer routing: route through v2 lowerer when FanOut/Join nodes are present (not just BoundedLoop), preventing silent flattening of parallelGroup structure in v1 pipelines without review-cycle. Commit reviewer's two test artifacts: - test/dogfood-full-feature.mjs (captures the B1 crash) - test/core/change-run/reviewer-workspace-lock.test.ts (proves lock invariant) Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ecp-full-feature): bind a synthetic capability for the FanOut evaluator The FanOut condition evaluator (and the v2-authored Choice evaluator) are produced by normalization, not by an authored stage, so no descriptor in the production capability catalog backs them. The profile resolver skipped them entirely, so `pipeline start full-feature` succeeded (plan creation does not consult the binding) and the Run then died at the first FanOut admission with: Error: No capability/policy binding for root:fanout:experts Task 3.4 called for a synthetic `parallel-dispatch` capability; it had never been implemented (zero occurrences in src/). Add it, plus `choice-select` for Choice, in BOTH the v1-migration and v2-authored resolvers, with matching policy stages. Bindings declare `workspace: { access: 'none' }` and no effects, matching the plan node the lowerer emits. analyzeReconcilerSupport must expect those nodeIds too: its shape check is a strict equality against the sealed capability list, and full-feature has BOTH a ReviewCycle BoundedLoop and a parallelGroup, so it takes the supported_v2_review_cycle branch - which also needed the FanOut evaluator in its expected set. Legacy `Choice` nodes from v1 `condition:` normalization carry legacyRuntimeOwner and are metadata carriers the lowerer skips; they must NOT get a binding, which is why the helper returns a nullable capability name rather than a type predicate. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ecp-full-feature): lower FanOut member paths as full plan paths createRuntimePlan resolves `fanOut.members[].nodeId` through the plan's path -> nodeId map, but the lowerer passed the v1 normalizer's unprefixed `stage:<id>` form. The lookup missed, the non-null assertion turned it into `nodeId: undefined`, and the reconciler's FanOut pass then found no matching atomic node for any member. The result was a SILENT stall: `pipeline status` reported all six members `ready` with `joinState: waiting` forever and nothing was ever admitted. Normalize member paths to `root:stage:<id>` once in the lowerer (the same string the member atomic node is lowered under), so the reconciler, the facade's FanOut completion validator, and the projector all match on one form. The FanOut condition result's `activeMembers` therefore names `root:stage:*`. Also make the unresolved case loud instead of silent: createRuntimePlan now rejects a fan-out or join node that references a member path with no plan node, rather than emitting `undefined` nodeIds. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ecp-full-feature): resolve the succeeded-set closure before candidate passes `reconcile` ran its succeeded-set contributions in a fixed pass order: bounded-loop (1b), choice (1c), fan-out (1d), join (1e). A node whose `requires` is satisfied only by a LATER pass therefore never became ready. full-feature's `review-loop` BoundedLoop requires the Join, so the bounded-loop pass always skipped it — and because `reconcile` is pure and recomputes the same order on every call, this was a permanent stall, not a one-tick delay. The real-CLI Run resolved the Join (`joinState: proceeding`, 6/6 members succeeded) and then sat with an empty frontier. Derive the transitive closure over the non-atomic node kinds up front, iterating to a fixed point, so readiness no longer depends on pass order. The `succeeded.add` calls in the passes below become idempotent no-ops and the emitted action order is unchanged. Factor the Join verdict into `evaluateJoin` (failed | waiting | proceed) and the loop terminal check into `boundedLoopIsClean`, shared by the closure and the candidate passes so the two can never disagree about when the barrier opens. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ecp-full-feature): size the sealed Record limits from the plan `counters.attempts` is a RUN-WIDE count of distinct attemptIds, so the flat `maxAttempts: 12` default capped the whole Run at 12 Actions regardless of `maxActions: 64`. full-feature needs 17 (3 lead-in + fan-out condition + 6 members + 4 review-cycle phases + ship/retain/archive), so the real-CLI Run escalated with `execution_budget_exhausted` ("Sealed attempts limit reached") in the middle of the review cycle — a Run that could never complete. Derive the ceiling from the plan: one Action per admittable node plus `maxIterations x body size` per bounded loop, doubled for retry headroom. Join and Finish nodes are excluded (a Join is never admitted). Limits only ever grow relative to the flat defaults; callers passing explicit limits are unaffected. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor(ecp-full-feature): type the parallel/choice section views (m3) `buildParallelSection` and `buildChoiceSection` returned `unknown | null`. The CLI `pipeline status` renderer, the Management API, and the Operations UI all read the SAME projection, so give it exported interfaces (ParallelSectionView / ChoiceSectionView, with ParallelMemberStatus and ParallelJoinState unions) instead. Also narrow two untyped wire-node fields in the Canvas FanOut/Join panels; `listValue` takes `readonly string[]` and was being handed `unknown`, which failed `packages/ui` tsc. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(ecp-full-feature): real-CLI full-feature dogfood + m2 support tests Replace the placeholder dogfood (which only captured the pre-fix start failure) with a driver that runs four REAL CLI scenarios end to end: A. parallel success - office-hours -> propose -> apply -> FanOut condition -> 6 members (cap 3) -> Join -> review-loop (findings round: review -> triage -> fix -> re-review) -> ship -> retain -> archive -> completed B1. optional member fails -> suppressed, Join proceeds to the review-loop B2. required member fails -> Join escalates, review-loop never starts C. restart idempotency - re-reconcile mid-FanOut from a fresh CLI process: same frontier, same action count, completed member not re-admitted Scenario A also captures `pipeline status` DURING the FanOut phase, which is the evidence task 13.5 asks for. m2: task 10.4 was ticked with no test. Add analyzeReconcilerSupport coverage - supported_v2_parallel for a parallelGroup pipeline with no ReviewCycle, reconciler available for full-feature's parallel+ReviewCycle shape, and a regression guard that a profile missing the synthetic FanOut binding is rejected as unsupported_pipeline_shape BEFORE a Run starts. Plus lowerer guards for the member-path bug: member paths must be full plan paths that resolve to the member atomic nodes, and an unresolvable member path must throw. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(ecp-full-feature): tick tasks 13.3-13.5 with real-CLI evidence Real CLI `full-feature` Run reached `completed`: RunId run:4805749044378972f4383d79b176a291c9b0a8c25b67e27edb405d5656a3b342 plan sha256:6a28a67428024221e720ea7613c642b857c8f77f766069c6e8973cd584a63a7d 17 Actions, recordVersion 34, engine=reconciler Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(ecp-full-feature): commit planner artifacts (proposal, design, specs, planning-context) The implementer committed tasks.md but left the planner's artifacts untracked. ship needs proposal.md for the PR body and archive needs the delta specs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ecp-full-feature): gate choice branch admission on the committed selection N1. The delta spec says "Un-selected branches SHALL never become eligible for admission". Execution violated that three ways, each proved by a reviewer probe (`test/core/change-run/reviewer-choice-semantics.test.ts`, all 3 failing at `c9459bd1`): 1. after the selected branch completed and the workspace lock freed, the ONLY admit emitted was the REJECTED branch; 2. no `finish` was ever emitted — the rejected branch held `finishCandidate` hostage, so a choice Run could only "complete" by also executing the branch it rejected; 3. a node requiring the SELECTED branch was admitted before that branch ran. Root cause of (3): `succeeded.add(branchNodeId)` on selection. The branch is an ordinary plan node; marking it succeeded because it was CHOSEN let everything downstream of it run before it had executed. It now earns succeeded status by committing its own Action, exactly like every other node. Root cause of (1) and (2): nothing ever excluded rejected branches. `choiceBranchExclusions` now computes, from committed Record truth, the nodes reachable ONLY from a rejected branch entry (transitively — excluding just the entry node would leave its downstream blocking the finish). Every candidate pass and `finishCandidate` skip that set. A convergence node reachable from BOTH the selected and a rejected entry is NOT excluded; it stays gated by its own `requires`. Also harden the kernel against a malformed selection: a committed result that does not name a DECLARED outcome no longer counts as a selection. Previously it still marked the choice succeeded, unblocking every branch at once; it now leaves the choice unresolved and re-admits the evaluator (bounded by the sealed attempt limit). Pre-existing in the original slice, but `23d47264` made it reachable: before the synthetic evaluator binding, v2-authored choice Runs died at the evaluator admission. `full-feature` does not exercise v2 Choice (its Choice nodes are legacy `condition:` metadata carriers the lowerer skips), so the ECP-4 exit evidence is unaffected. The reviewer's 3 probes are kept verbatim as the regression tests. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(ecp-full-feature): facade choice/fan-out rejection tests + close the validator bypass N2. Task 7.4 was ticked but the tests existed NOWHERE — zero grep hits for the validators' error strings across `test/`. `facade-evaluator-validation.test.ts` now drives the real facade `complete()` path (verifyCompletion → validators → commit) for both evaluators: valid results accepted, undeclared outcome rejected, missing `activeMembers` rejected, and a condition that suppresses a REQUIRED member rejected. The gap also hid a bypass. Both validators used "result is a non-array object" as the PRECONDITION for validating at all, so a string result skipped validation entirely, committed, and left the Run permanently stalled with no branch selected and no diagnostic. A non-object result is now rejected with a message naming the shape received. Scope validation to `status: 'succeeded'` while tightening it: a failed or blocked evaluator legitimately carries no selection, and demanding one would make the failure unrecordable. Two tests pin that. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(ecp-full-feature): parallel/choice section projection + cross-plane parity N3. Tasks 8.5/8.6 were ticked but no test anywhere asserted either section -- projector.test.ts, cross-plane-parity.test.ts, view-invariants.test.ts and the management-api suite all returned zero hits for "parallel", and choice/1 had no coverage of any kind. The only evidence was the CLI-side dogfood, which covers one plane and whose pipeline has no v2 choice node. Adds unit coverage for member status derivation (waiting / suppressed / ready / running / succeeded / failed), joinState transitions (not-reached / waiting / proceeding / failed), budget usage, keyBlockers, and optional-failure suppression; plus choice outcome and active-branch marking. Parity: one Record asserted through the projector, the real CLI `pipeline status --json`, and `handleRunDetail` off a filesystem store -- all three must see byte-identical sections. Task 8.6 also names Operations, now annotated in tasks.md as NOT covered: packages/ui has no consumer of these run-view sections yet, so there is nothing to assert parity against. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ecp-full-feature): enforce the binding-shape check on the parallel branch N4. Task 10.3 promises `supported_v2_parallel` "when all FanOut/Join/Choice bindings are present", but the branch returned supported whenever ANY profile existed -- unlike the v2-executable and v2-review-cycle branches it never compared the sealed capability list against the expected nodeId set. A parallel-only pipeline with an incomplete binding set therefore reported supported and then died mid-Run at admission: exactly the failure mode `23d47264` fixed on the review-cycle branch. Apply the same strict comparison (root AtomicStages plus orchestration evaluators). KNOWN ADJACENT GAP, deliberately not fixed here: for a v1 pipeline with parallelGroup but NO ReviewCycle loop, `resolveCapabilityBindings` still takes the plain v1 path and emits `stage:<id>` bindings, while the v2 lowerer that `requiresV2Lowering` routes it through looks up `root:stage:<id>`. Such a pipeline cannot lower today regardless of this check; the check makes that visible before the Run starts instead of crashing at plan creation. Widening the resolver is a real capability addition with no built-in consumer to dogfood against, so it is reported rather than shipped unverified. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(ecp-full-feature): make the dogfood gate on the behaviour it prints N5. `scenarioErrors` counted thrown exceptions only, so a regression that made B1's `joinProceeded` false or B2's `reviewLoopStarted` true would still exit 0 -- the FINAL flags were decoration. As the durable 13.3-13.5 regression net the script now checks 11 behavioural assertions (terminal outcome, Join resolution counts, the parallel section and its member frontier during FanOut, the concurrency cap, the full review-loop phase sequence, optional suppression, required-failure escalation, restart stability) and exits non-zero on any violation. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(ecp-full-feature): align the delta spec with the executed choice/join semantics Two requirements described behaviour the kernel does not (and should not) have. Choice: the spec said the reconciler "SHALL add branches['complex'] to the succeeded set" on selection. Taken literally that is what caused N1's third violation -- marking a branch succeeded because it was CHOSEN let everything downstream of it run before the branch had executed. Selection makes the branch ELIGIBLE; the branch enters the succeeded set by committing its own Action. Adds scenarios for the rejected branch staying ineligible after the selected branch completes, the rejected branch not blocking the implicit finish, a convergence node reachable from both branches not being excluded, and a malformed result not counting as a selection. Join: `evaluateJoin` is stricter than the inlined logic it replaced. The old optional-member loop counted only `state === null` as non-terminal, so a Join could proceed while an optional member was still in flight. The new code waits, which is what the requirement text and task 6.3 always intended. That behaviour is now load-bearing, so state it explicitly and add a scenario for it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ecp-full-feature): a FAILED evaluator result is not a decision R2-M2. `committedResultForNode` was status-blind, so a FAILED orchestration evaluator still drove execution. Proved by 2 reviewer probes (`test/core/change-run/reviewer-r2-evaluator-status.test.ts`), both failing against `f1951b52`: - a choice evaluator that crashed mid-analysis but left `{ outcome: 'simple', error: ... }` SELECTED the branch and admitted it; - a fan-out condition that crashed with `{ error: ... }` dispatched ALL members, because `readActiveMembers` treats a result with no `activeMembers` array as "every member is active". Renamed to `succeededResultForNode` and filtered on the committed status, so every call site reads as what it means. Only a succeeded completion is a decision; anything else leaves the evaluator unresolved and re-admitted. Pre-existing kernel behaviour, but this delta touched the seam twice without closing it — `12cc6131` scoped the facade validators to succeeded completions, which formalized the door this leaves open on the kernel side. Both reviewer probes are kept verbatim as regression tests. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(ecp-full-feature): keep the r2 reviewer evaluator-status probes as regression tests Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ecp-full-feature): reject plans whose rejoin depends on both choice branches R2-M1. The round-2 carve-out ("a convergence node reachable from BOTH branches is not excluded; it stays gated by its own requires") was accurate to the code and concealed a silent deadlock. `requires` is AND semantics and exactly one branch is ever selected, so a rejoin node downstream of both can never have all its requires satisfied. Not excluded means never admitted; still counted by the implicit-finish check means the Run stalls at `waiting` forever, with no escalate and no diagnostic. My own round-2 test asserted `finish: false` as the EXPECTED outcome, which is how the gap survived review. Rejected at `createRuntimePlan` rather than escalated at runtime. This is a purely static property of the plan — no Record is involved and the verdict cannot change as a Run progresses — so the build-time check fails `pipeline start` outright with the offending node and choice named, instead of creating a Run that stalls and only later escalates. Runtime detection would re-derive the same static fact on every reconcile and surface it after the fact. Same reasoning as rejecting unresolvable fan-out/join member paths in `3f4b0e53`: loud at build time beats silent at runtime. Detection is transitive (a rejoin two hops below each branch is caught) and scoped per choice, so branches of DIFFERENT choices may still converge — that shape is satisfiable when both choices pick the converging outcomes. Replaces the round-2 convergence test with three: direct rejoin rejected, transitive rejoin rejected, and non-converging branches still accepted. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs(ecp-full-feature): spec the evaluator status rule, retry cap, and rejoin rejection Three corrections to the delta spec, matching this round's kernel changes. R2-M2: the rewritten Choice requirement was status-silent — it said a committed result naming a declared outcome counts as a selection, without requiring the completion to have SUCCEEDED. A failed evaluator can carry partial output naming an outcome. Both the Choice and FanOut requirements now say only a succeeded completion is a decision, with scenarios for each. R2-m1: the unresolved-evaluator re-admission had no stated bound. Both requirements now say re-admission is bounded and the reconciler escalates with a code naming the evaluator, with a scenario each. R2-M1: the requirement claimed un-selected branches "SHALL NOT block the implicit finish", which was transitively violated by a rejoin depending on both branches. It now states that such a plan is rejected at plan creation, and the "convergence node is not excluded" scenario is replaced by a rejection scenario plus one confirming that non-converging branches are still accepted. Follow-up recorded for the ticket, deliberately NOT fixed here (out of ECP-4 scope per the round-2 review): `supported_v2_parallel` is production-unreachable until `resolveCapabilityBindings` emits `root:`-path bindings for parallel-only v1 shapes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-full-feature): consume parallel/1 + choice/1 in the Operations plane proposal.md promises the parallel frontier is "visible from CLI, Management API, and Operations". The first two planes read the projector's `parallel/1` and `choice/1` sections; `packages/ui` had no consumer of either -- its only section extractors were `getRootDagSection` and `getReviewCycleSection`. Mirror the core wire types (`ParallelSectionView`, `ParallelMemberView`, `ParallelMemberStatus`, `ParallelJoinState`, `ChoiceSectionView`, `ChoiceBranchView` from `src/core/change-run/internal/projector.ts`) into the UI types module, add `getParallelSection`/`getChoiceSection` alongside the existing extractors, and render both sections in the Run detail body: fan-out members with status/role/condition, the Join barrier state, budget usage, the projected member frontier, and the choice outcome with each branch marked active or inactive. `joinPath` and `outcome` are mirrored as OPTIONAL properties: the server type is `string | undefined` and JSON drops undefined keys on the wire. Every value is read from the projection. The UI recomputes no member status, join state, count, or branch activation -- the same discipline it already follows for the root-DAG frontier. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(ecp-full-feature): Operations-plane parity for parallel/choice (8.6) Task 8.6 asks for parity across CLI, Management API and Operations. The first two were covered by `test/core/change-run/projector-parallel-choice.test.ts`; the task was ticked with an annotation admitting the UI plane was missing. Add the fourth plane under `packages/ui/test/`, per the existing `cross-plane-parity.test.tsx` convention: render the REAL `OperationsSection` against the SAME fixture the node-side parity test builds (FanOut `root:experts` cap 2 / budget 3 with the condition committed and the required member settled; Choice `root:pick` committed `simple`) and assert the DOM carries the projected member statuses, roles, conditions, join state, budget usage, frontier counts, key blockers, outcome and branch activation. The canonical section constants are the verbatim `projectRunView` output for those Records, not values written to match the component; their provenance is documented in the file header. Three assertions feed deliberately INCOHERENT sections -- counts that no tally of `members` produces, a `waiting` join over all-succeeded members, and an active branch that is not the selected outcome. With coherent fixtures a UI that recomputed these is indistinguishable from one that consumes them; the disagreement is what pins the behaviour. Verified by mutation: recomputing `activeCount` from member statuses and branch activation from `outcome` turns those tests red. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ecp-full-feature): projector must read the LAST SUCCEEDED evaluator result R3-M1. The `succeededResultForNode` sweep closed the kernel but missed the projector, which had a second defect on top: it took the FIRST result-bearing attempt where the kernel takes the LAST, and ignored completion status entirely. Both defects only matter together, and the retry cap is what made them live. After a failed-then-retried-successfully evaluator — now an ORDINARY path precisely because retries exist — the projection kept reading the failed first attempt forever: members `waiting`, joinState `not-reached`, while the Run was actually executing. It also displayed a crashed evaluator's partial output as a real selection. Same shape as N1: a fix makes a dead path live, and the newly live path carries its own bug. `succeededEvaluatorResult` mirrors the kernel reader — last attempt, succeeded only, object-shaped — and both section builders use it. Display-only; execution semantics are untouched. All three parity planes share this one reader, which is exactly why the parity suite could not catch it: every plane was wrong identically. Reviewer probes in `reviewer-r3-ui-constants-provenance.test.ts` kept verbatim. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ecp-full-feature): evaluator retry cap must not count the in-flight attempt R3-m1. `occurrenceForNodeId` counts EVERY invocation including one still active, so with 2 failures committed and the third attempt in flight, the next reconcile — any settle, a `resume-run`, or another node completing — escalated `choice_evaluator_unresolved` and terminated a Run that was about to resolve, discarding live paid work. `evaluatorAttempts` reports both facts separately. While an attempt is active no verdict is due at all: the pass emits neither an admit (which would double-dispatch) nor an escalate. Only committed attempts count against the cap. The round-3 cap tests passed honestly because their records are all-terminal; they still do. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(ecp-full-feature): keep the r3 reviewer probes as regression tests 3 probes that proved R3-M1/R3-m1 plus 4 provenance tripwires pinning the UI parity constants to the real projection. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(ecp-product-closure): Operations consumes review-cycle/1 + kernel-anchored parity constants ECP-1's shipped `executable-review-cycle` delta requires CLI, Management API and Operations to consume the same ChangeRunView review-cycle section. Two of the three planes did; `getReviewCycleSection` had zero consumers in `packages/ui/src`. Adds a `ReviewCycleSection` renderer to OperationsSection following ECP-4's parallel/choice pattern: round against cap, phase, outcome, findings with severity + status, bound actors, wait reason - every value read from the projection, none re-derived (the round cap, the cle…
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
All context-occupancy thresholds (handoff default 0.5, reuse default 0.25, per-role overrides) are fractions of the context window. For 1M-context models a fraction works, but for small-window models (e.g. the GPT-5.6 family at ~300K) a fraction fires far too early in absolute-token terms — and those models degrade gracefully near their limit, so the remaining absolute headroom is the better signal. Configs need an absolute-token threshold form, and the CLI should ship sensible per-model presets so users of small-window models get good behavior without hand-tuning.
What Changes
thresholdvalues (pipeline-level, per-role, stage-level) accept a second form:{ remainingTokens: <positive integer> }— an absolute required-headroom threshold — alongside the existing bare fraction in (0, 1]. A bare number is ALWAYS a fraction; the absolute form is ALWAYS the object, so no value is ambiguous.src/core/model-presets.ts) keyed by model-id substring patterns provides, per known model family: the context-window size and suggested handoff/reuse threshold values. It subsumes the existing ad-hoc model→limit map inresolveModelLimit.handoff> pipelinehandoff.roles[<role>]> pipelinehandoff> model preset (via the stage's resolved model) > built-in default. Reuse:reuse.roles[<role>]>reuse.threshold> model preset (viaagents[<role>]model) > built-in default. Any ordinary config value therefore overrides a preset.rasen pipeline show --jsonreports resolved thresholds in whichever form they resolved to (bare number for fractions — byte-identical for existing configs — or the{ remainingTokens }object), andsourcecan now bepreset.rasen agent contextoutput gainsremainingTokens(limit − contextTokens) so absolute thresholds can be compared directly against a probe.pct >= thands off /pct <= tallows reuse; absolute:remainingTokens <= Nhands off /remainingTokens >= Nallows reuse).Review
Went through 3 rounds of an independent review-cycle loop — clean, with all Major findings resolved (test-tautology fixes for preset ordering and byte-identity guarantees, and a
sourceprovenance fix sopipeline shownever mislabels a preset-sourced threshold as coming from a config layer). Full findings in the change's review-report.md.Test plan
dev/0.1.4's concurrentfix/codex-host-compatmerge and rebuildingdist/)tsc --noEmitcleanrasen validate handoff-absolute-thresholds— validhandoff: { threshold: { remainingTokens: 60000 } }validates and shows correctly; a fixture withagents.implementer.model: gpt-5.xand no thresholds resolvessource: preset🤖 Generated with Claude Code
https://claude.ai/code/session_011an7Av8HWb8tWFrUPD9ZTi