raghav and aaron reconciliation branch - #90
Merged
Conversation
Scoped to the pulse-designer persona (never hijacks the default build agent),
one-question-at-a-time, 8 stages PLATFORM→CALIBRATE, transmon fully wired /
Rydberg recorded-honestly, amicode_* tools as bookkeeping with bash+amico-run
still the launch mechanism. Tests: interview block + no-unknown-placeholder
assertion (only {{TEMPLATE_PATH}}/{{JULIA_PROJECT}} substituted).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…gin load / live turns) Tier A green vs stock 1.17.3; B gates on the plugin file, C on creds (none on this machine tonight — morning step: opencode auth login). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Four bookkeeping tools (pick_system / set_model / formulate / solve-record) executed by opencode's embedded Bun runtime; entities as TOML under ~/.amico/runs/default/_entities ($AMICODE_ENTITIES_DIR override). Registration via buildOpencodeConfigContent: plugin abs-path + pulse-designer agent block + entities-dir permission grant (T8 probe: both mechanisms live at 1.17.3). Plugin stays outside the extension bundle/tsconfig; load line on stderr (debug-config stdout purity — regression-guarded). 99→122 tests green. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Free anonymous provider (opencode/big-pickle) resolves without auth.json → AMICODE_E2E_LIVE=1 forces the creds gate. Cadence assertion = stage-batching check (multiple '?' inside one platform question is fine); transcript saved to tmpdir for the experiment note. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Real-home mode now substitutes AGENTS.md via prepareOpencodeProject (stage 6 needs real template/julia paths). Keyword-routed answers, bounded turns, run-dir poll. Result: 10 turns on the free model → all four amicode_* tools fired (System→Formulation→Run stub) → authored solve.jl → amico-run launch → F=0.9998977/60 iters. AMICODE_E2E_FULLCHAIN=1 gates it. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
amicode_to_hardware (device_session stub; gate + checks PINNED by the serializer so a stub can never claim an approval that didn't happen) and amicode_calibrate (ILC stub, auto-refs device_session.toml). No device I/O — explanations say so explicitly, mirroring AGENTS.md stage 8, which now names both tools. 6 tools registered (verified via /experimental/tool/ids). 122→128 tests green. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
'Who are you?' now answers 'I'm Amico — Amicode's pulse-design copilot' (never 'opencode, an interactive CLI tool'), and a greeting/no-intent session opener triggers the stage-1 PLATFORM question instead of a generic assistant reply. Specific asks keep the skip rule. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Interview option-stages (platform, sim-vs-solve, problem/gate) now go through amicode_ask(question, options[2-6]); the renderer draws buttons from the tool part's input and a click sends the option text as the next user message. Free-form values stay plain questions. 7 tools registered (runtime-verified). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Live-demo findings: the free model repeated the question in prose and answered it itself. Tool return now says STOP HERE explicitly; AGENTS.md mirrors it. New optional details[] arg — one dim qualifier per button (minimalist card, slightly more context). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Pixel <0110> face in brand yellow on the H silhouette (from the website robot); currentColor body serves dark+light. Same asset ships in the fork's logo/favicon (kept in sync manually — canonical geometry recorded there in AMICODE-PATCHES.md). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ited stop Work found staged in the night tree (parallel session / Aaron) — folded in after validating build + full suite green. Fixes the Remote-SSH pain: forward 43117 once, restarts reuse it; stop() now resolves on child exit so a fixed-port restart can't race the old process for the socket. Set 0 to restore per-start ephemeral ports.
…tlement registry)
…N, plugin-sidecar pattern)
…e_manifest transport, scores-root grant, never-brick fallback
…s (score_guard sibling module, entitiesDir transport)
…ded-state (v2) The compiled score is now what runtime users see — the STOP-HERE ask discipline and optional per-option details move in from the AGENTS.md fallback, plus the anti-drift anchor from Aaron's live session (rail said rydberg while the interview asked transmon-omega): re-read the System entity before stage-2+ questions; correct the record first if it's wrong. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ion cross-check Score content bumps (v1->v2 tonight) must not red the suite; instead assert the compiled banner and the manifest agree on whatever version ships.
… deprecated Live finding: 1.17.3 ships a turn-blocking question tool (tool/question.ts) with options+descriptions+custom answers — the model reached for it unprompted and its semantics structurally kill self-answering and prose repeats. Score #0 v3 + AGENTS.md fallback now route all choice questions through it; amicode_ask stays registered but deprecated. One mechanism, upstream-owned. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ine talk The starter chip hit a Webfetch to the engine's website and a flat 3-sentence reply. The canonical question now has a canonical answer: capability bullets (interview, fast-path, live inspector, warm-start/resume, hardware preview), honest scope line, then next-step options via the question tool. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…late fetch_opencode gains repo/tag manifest fields with authenticated gh download for the private harmoniqs/opencode release (206 tests). TESTING.md = the 7-point branch test script for the team. solve_rydberg_cz.jl: 2-atom 3-level QuEra deep-blockade CZ, public-Piccolo-only (fixed-phase NLP + post-hoc virtual-Z scan; free_phase kwarg is Piccolissimo-gated) — vetting solve in flight, score wiring lands with its result. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Point an agent at the repo and it can set up, verify, and test the branch: prerequisites with checks, ordered setup, five verification gates, dev facts (fixed port, fork boundary, plugin runtime, scores), sharp edges.
…icode.1) Lock repointed to the mirror (both platform shas from the release assets); clean-fetch verified end-to-end via the authenticated gh path — downloaded binary is byte-identical to the deployed build (8aa718ad…) and serves <title>Amicode</title>. Rydberg CZ template clearly marked EXPERIMENTAL (vetting stalled on pathological first-iteration NLP cost; parked) and TESTING.md de-overclaimed accordingly. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…cker advertised a wedge as '$(pulse) live', contradicting the status bar Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…oard replies gated on panel visibility — a hidden panel rendering LLM-driven content must not sample the OS clipboard in the background Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… 'v1' while SCORE.md is v3, red on every creds-bearing machine Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…'s) and the one-spine mirror comment now names the stall threshold + display vocabulary as mirrored semantics Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…'s export); one-spine mirror comment names the stall threshold + display vocabulary as mirrored semantics Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ettierrc (printWidth 120, semi) — the repo had no formatter configured; config formalizes the existing conventions Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…the previous style commit) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…e full-repo normalization so the formatter is enforceable from here Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…t resolution gambles on provider ordering and (with Google creds) picked a hanging preview model for every headless/agent turn; anthropic > GA gemini flash, and a user's global model always wins (1.17.3 preserve contract) + tests Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… the amicode_ask TOOL CALL and no prose, so the ask input (question+options) counts as the turn text; production model pin wired into the e2e servers Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…es through a save dialog (downloads are dead in the framed app); PNG-only, basename-sanitized, size-bounded, relay-allowlisted Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…is capacity-throttled at peak ('model overloaded' → failed turns render as 'model undefined' stubs in chat); the GA flash answered in 1.4s during the same window
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ee (zen free tier, no user quota; answered a tool-bearing turn in ~3s while Gemini kept capacity-throttling); anthropic still wins when creds exist, user global model always wins Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
fast/vsix-gate/boot-smoke fetch the vendored opencode binary from the PRIVATE harmoniqs/opencode release via `gh release download`. The default Actions GITHUB_TOKEN is scoped to this repo only, so gh is unauthenticated for harmoniqs/opencode and the fetch exits 1 — reding every job that vendors opencode (schema-roundtrip is untouched; it needs no binary). Pass a cross-repo token as GH_TOKEN to the three fetch:opencode step defs (boot-smoke's single step covers both matrix legs). Mirrors #80, adapted to this branch's split vsix-gate step. Requires repo secret OPENCODE_FETCH_TOKEN: a fine-grained PAT scoped to harmoniqs/opencode with Contents:Read (validated end-to-end — downloads both assets and the bytes match opencode.lock.json's SHA256 gate). Stays red until the secret exists; green once it's set, no further push needed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Pin the vendored opencode binary to the new fork release cut from rchari/amicode-fixes (opencode PR #1): Gemini tool-schema fix, /amicode run-cards + profile endpoints, one-spine run truth, multi-drive pulse fix, save-file bridge. darwin-arm64 + linux-x64 sha256 updated. vsix build verified locally (linux-x64 fetch + sha match + vsce package -> amicode.vsix). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…/stop truth, model pin, audit sweep, formatter (#89) * feat(1.1): Scheduler — serial run queue built to the ratified Executor contract (#56) `enqueue(spec, {concurrent?}) → ScheduledRun{queueId, handle: Promise<RunHandle>, cancel}` plus a multi-consumer lifecycle stream (queued/started/finished/cancelled/error) for RunsManager/StatusBar (1.2). Lives in @amicode/amico-run (node-only, no vscode) beside the LocalExecutor it drives. Built TO the Track C contract (ratified 2026-07-02) so Δ8's RemoteExecutor drops in with zero reshape: - S12: enqueue resolves to the executor's RunHandle UNTOUCHED (identity passthrough, pinned by test) — downstream never sees an executor type. - (b) abort() is a request, not a kill: the pump advances ONLY when `finished` resolves; a post-abort() run still holds the queue (pinned by test). - (c) per-executor warming budget: the Scheduler owns NO timers — structurally pinned (test greps the source for setTimeout/setInterval). - (d) `finished` never rejects per contract; a rogue rejection is survived (error event) rather than wedging every queued run. Semantics: strictly serial; cancel() dequeues only pre-start (a live run is stopped via RunHandle.abort(), never the queue); a submit() ConfigError rejects that entry's handle, emits `error`, and the queue advances; `concurrent: true` is the NAMED Phase-4 seam — rejected loudly (ConfigError) instead of silently serializing. Listener errors are isolated from the pump. TDD: 12 tests (RED first) — serial ordering, S12 identity, opts passthrough, abort≠terminated, lifecycle sequence with positions, cancel pre/post start, ConfigError advance, the concurrent seam, multi-listener + dispose, throwing listener, and the no-timers structural pin. Repo suite green (schema 34, amico-run 59, extension 80); build + typecheck clean. Closes #56. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(1.1): scheduler review — microtask-deferred re-pump + enforce the defensive claims Adversarial review (mutation-tested; 0 must-fix, 2 should-fix) — all folded in: - finally's re-pump is now queueMicrotask-deferred: a contract-violating executor whose submit() throws SYNCHRONOUSLY previously made the finally a direct recursion — a backlog of such failures accumulated behind a pending run blew the stack on drain (RangeError) and STRANDED the rest of the queue. Mutation-verified: reverting to the direct call fails the new test (RangeError + timeout); the deferral drains 8000 sync-throwers flat. - The rogue-`finished`-rejection branch is now enforced, not just advertised: new test pins error-event + queue-advance (mutation-verified: deleting the branch fails it). Rejection reason normalized (instanceof Error) to match the submit path. - Explicit unhandledRejection pin for the internal handle.catch suppression (an untouched ScheduledRun.handle never trips the process on cancel). - cancel() docstring now names all three false cases (started / already cancelled / mid-submit, where handle may still reject); no-timers grep widened to setImmediate|Date.now. 15 tests green; repo suite green; typecheck clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: catalog entry card — components, save-to-catalog flow, session catalog (UX2) (#73) * feat(47): catalog entry card — components, save-to-catalog flow, session catalog (UX2) The catalog card (Krishna p5, UX2) shipped end to end as a seam prototype. Identity is Hamiltonian-anchored and user-named; the card is the feedback artifact for the open field-selection questions. Components (media/ui): - chip: the entry's handle — gate · system · tags · #index. The system slot carries the USER-ASSIGNED name (researchers think in named devices, "Emerald-Q3"); derived family is the fallback. Tags render as dashed "proposed" segments — none of tags/index/name are in catalog-entry.schema.json, and the marking is deliberate. - catalogcard: header (chip + run-id + Tune/Warm-Start/Promote actions, VS Code secondary-button styling, right-justified) → pulseplot (reused, hydrated) → metadata / pulse-data / high-level-metrics panels → sibling-chips row. Metrics are uniform hero cards in a wrapping flex row: fidelity · gate time (params.T) · spectral bandwidth (95%-power, computed client-side from the knots, DC removed; definitional choices marked proposed) · robustness (empty proposed slot — needs perturbed rollouts recorded at solve time). Solver telemetry (iterations/wall) lives in metadata: provenance, not pulse quality. Drive labels are u_i, matching the trajectory component and Piccolo's plot defaults. Flow + shell (src): - Save to catalog: a converged run's promote prompt (live) or the demo replay prompt → optional system-name and tags input → a pointer entry in the session catalog (workspaceState — NOT the Phase-3 CatalogStore; Q91/Q92 open) → the card opens, hydrated from the real run artifacts (run.toml identity; result.toml fidelity/params with gate/system lifted; run.log pulse lines → the plot). - Catalog tree view: chip-shaped rows (gate · system · fidelity, tags in description/tooltip), newest-first, deduped by run_id; click reopens the card. amicode.catalog.refresh actually registered. - The solve-template change that records params.gate/params.system on NEW runs ships separately; the bundled demo fixture already carries them (schema-validated). Tests: pulse-line grammar fixtures, hydration (field mapping, gate lift, newest-record pulse, degradation), session tree (dedup, order, open-card command, empty state). Closes #47's Kate-lane scope as amended (actions row instead of the "what do next" section; resume hand-off design returns to UX1 — see the deferred-items comment on #46). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(47): remove-from-catalog — non-destructive unsave via the tree context menu Right-click a Catalog row → "Remove from Catalog": deletes the POINTER record only (workspaceState) — the run dir and pulse.jld2 stay on disk, per the raw-data-trust principle. Archive/supersede lifecycle belongs to the Phase-3 CatalogStore (Q94/Q95), deliberately not faked here. Kept off the command palette (when: false) — the action only means something on a row. Tests: removal preserves remaining order, is idempotent, and rows carry the context value gating the menu. Part of #47. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(47): reveal-or-create dedupe for catalog-card panels (review) Clicking the same catalog row repeatedly spawned duplicate panels. Per Jack's review: run_id → live-panel map, second open re-focuses via reveal, disposal cleans the map so a closed tab re-creates fresh. Shell test covers all three; vscode stub grows createWebviewPanel + registerCommand to support it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * 1.2 — RunsManager: multi-run engine keyed on the append-only runs/index (#70) * feat(1.2): RunsManager — multi-run engine keyed on the append-only runs/index (#57) Replaces β's single-run RunsRootWatcher. Discovery now tails `runs/index` (amico-run's appendIndex TSV) instead of following the `latest` symlink — `latest` keeps being written (frozen contract) but a second concurrent solve no longer yanks tracking off the first mid-flight: every run WITHOUT a FINISHED gets its own pipeline (replay → run-dir watch → run.log tail) and is tracked to completion. - run_registry.ts (pure, vscode-free): parseIndexLine (tolerant of torn/blank lines — the tail heals) + RunRegistry (idempotent by runId; iter high-water; FINISHED-keyed terminal state). - log_tailer.ts: LogTailer extracted verbatim from file_watcher.ts — reused for every run.log AND the index (both append-only). - runs_manager.ts: per-run pipelines with per-run PulseStream/SinkDedup; runId-gated routing — the SELECTED run drives the single-run Inspector + StatusBar (selection auto-follows the newest started run, β latest-follow parity; `selectRun` is 1.3's seam), while completions + the promote-once prompt fire for EVERY run, selected or not. Completion keys on FINISHED (never result.toml presence). Poll backstop + idempotent consumers as before. attachScheduler consumes the #56 lifecycle (structural SchedulerLike so this is independent of #68's merge): `started` registers + selects immediately. - extension.ts: RunsManager replaces the watcher; replayDemo now renders via EXPLICIT selection (pokeDiscovery + selectRun) since a finished-at-discovery run registers quietly (idle-at-launch parity) — promote stays suppressed. - file_watcher.ts deleted (superseded); its statemachine tests ported to runs_manager.test.ts (idle-on-finished, warming→FINISHED-keyed completion, #66 pulse tail routing, replay-seeded meta) + new multi-run coverage: concurrent runs both tracked with background completion/promote, re-select replay without promote re-pop, missing-dir index lines tolerated, the Scheduler seam, and the demo-replay explicit-selection path. 15 new tests; repo green (extension 103, amico-run 47, schema 34); typecheck + build clean. Closes #57. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(1.2): runs-manager review — disk-checked warming, scheduler metadata backfill, watch guards Adversarial review (0 must-fix, 2 should-fix) — applied: - selectRun re-checks DISK for FINISHED before posting warming (β parity): registry phase can be ≤700ms stale, and warming-after-completion inverted the terminal badge — real once 1.3's user-driven selectRun lands. - RunRegistry.backfill: a scheduler-registered run (runId+runDir only) gains createdAt/scriptPath when its index line lands — the 1.3 trees would otherwise see undefined metadata on every scheduler-launched run. - fs.watch 'error' listeners on all three watcher sites (root, per-run dir, tailer) — an unhandled FSWatcher error is an uncaught host exception; the poll backstop keeps things live. - RunRegistry.all() returns copies (1.3 callers can't mutate registry state); honest header comment on the transient replay/tail re-delivery window. - test/log_tailer.test.ts (review gap — the tailer is now load-bearing for discovery): torn-line carry-over, truncation re-read, startOffset contract, poke self-attach. 19 tests green, typecheck clean. Still owed next session (review nits, test-only): cross-run PULSE routing gate test, backfill unit test, stale-warming pin test; inspector pendingPulse reset on selection switch is deferred to 1.3. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(1.2): close the owed #70 review-nit gaps (all mutation-verified) The three test gaps the RunsManager review flagged: - RunRegistry.backfill: fills ONLY missing metadata (first-registration wins), no-throw on unknown runId — mutation-verified (dropping the missing-only guard fails it). - RunRegistry.all(): returns copies — a mutated snapshot can't corrupt registry state. - cross-run PULSE routing: a background run's pulse RECORD is gated on selection (not just iter) — never reaches the inspector while another run is selected. - stale-warming: selecting a run whose FINISHED landed inside the ≤700ms poll window shows completion, NOT warming (no terminal-badge inversion) — mutation-verified (reverting the disk re-check fails it). extension 111 tests (+8), repo green (amico-run 47, schema 34); typecheck clean. * fix(1.2): review #70 — pinned selection, torn-FINISHED retry, markFinished guard, single-pass discovery, one terminal orchestration Addresses jack-champagne's static/design pass, one commit per nothing — all five findings land together because #1/#4 reshape the same registerRun path: #1 (design, the #72 seam): explicit selectRun PINS the selection; auto-follow (newest registered live run, β latest-follow parity) only applies while nothing is pinned. A background solve starting can no longer yank the view off a run the user deliberately opened. Mutation-verified. #2: a FINISHED that is present but torn/invalid at discovery no longer finalizes as status:undefined-forever — the run falls through to the live path, whose checkFinished re-reads next tick (the retry the live lane already had). Promote stays suppressed (launch replay). No warming either (disk-checked: FINISHED exists). #3: RunRegistry.markFinished guards phase itself (first terminal wins) — a stray re-mark can't leave status:"failed" beside a stale fidelity. The guard now lives on the public surface, not only in completeRun. #4: discovery ingests the run dir ONCE. Auto-follow assigns selection BEFORE the registration replay, so the single pipelineSink pass both seeds state and feeds the display through routeIter/routePulse's selection gate — displaySink now backs only explicit selection replays. #5: the FINISHED→status→result.toml→fidelity orchestration lives ONCE, in run_dir_reader.readTerminalState (say-why callback preserved via the manager's channel); ingestRunDir and RunsManager.readTerminal both delegate, so a contract change (e.g. #64's formulation.toml) is edited in one place. Also rebased onto main (#68 Scheduler, #77 vsix-gate, #79 permission grant). 118 extension tests pass; typecheck + build clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * 1.3: Run Inspector single→multi-run (runId-keyed protocol, per-run panes) (#72) * feat(1.3): Run Inspector single→multi-run (runId-keyed protocol, per-run panes) Freeze-2 reshape (#58): the host↔webview message protocol is now runId-keyed and both the host and the webview fan into per-run panes. Single→multi only — the pane markup stays the current pulseplot (design lane, UX4 #49). Host (run_inspector.ts): a PaneBuffer per runId + activeRunId; runId-keyed surface postPulse/postIterationRecord/postCompletion/setWarmingUp/setRunLabel + new activate(runId). Per-run 5 Hz pulse throttle. resolveWebviewView replays EVERY pane from its buffer (S36) with positional ordering, then posts activate last. setWarmingUp guarded from clobbering a pane that already has data/terminal state. pulse stays plot-only (deliberately does not clear warming). Webview (media/ui/views/inspector.ts): createPanel() instances the former single-run view per runId (no shared globals); a router keys panels by runId, activate toggles the one visible pane, background/late messages only touch their own pane. Pane-hiding uses two-class selectors so it wins over layout.css `.stack` on specificity, not stylesheet order. RunsManager (runs_manager.ts): fans every run's live events into the inspector runId-tagged (routeIter/routePulse ungated); registration replay is state-only so the selected run never double-posts; selectRun adds activate; the single status bar stays selection-gated; completion + promote still fire per-run. Tests: runs_manager + inspector_view_contract updated to the runId-keyed API and fan-out semantics; added per-run isolation, per-run-throttle independence, S36 reopen, activate-last, warming-guard. New happy-dom webview test covers the router itself (per-run isolation, activate toggle, empty-state, plot-only pulse) — closes the coverage gap flagged in adversarial review. All invariants mutation-verified. 121 tests pass; typecheck + build clean. S6 (formulation preview) deferred — no formulation-emit in the frozen contract. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor(1.3): flow RunCompletion whole through completeRun — the #84/#81 seam Jack's #72 merge-seam heads-up: #81 adds `formulation?` to RunCompletion, and completeRun was the third completion path cherry-picking fields positionally (runId/status/fidelity) — once #81 landed, live-completed runs would carry formulation: undefined while replayed runs got it (the exact bug Kate caught on onFinished, reintroduced here). completeRun now takes the WHOLE RunCompletion; both feeders (ingestRunDir's sink verbatim, checkFinished via {runId, runDir, ...readTerminalState()}) funnel the object from the one shared read. An additive field is now a one-place edit (RunCompletion + readTerminalState) and reaches every consumer by construction — consumers cherry-pick at the leaf. Documented as the #84 funnel on both the type and completeRun; the full N-reader consolidation (catalog hydrator etc.) stays #84. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * 1.4a: smoke corpus — Scheduler → executor → RunsManager → Inspector e2e (#78) * feat(1.3): Run Inspector single→multi-run (runId-keyed protocol, per-run panes) Freeze-2 reshape (#58): the host↔webview message protocol is now runId-keyed and both the host and the webview fan into per-run panes. Single→multi only — the pane markup stays the current pulseplot (design lane, UX4 #49). Host (run_inspector.ts): a PaneBuffer per runId + activeRunId; runId-keyed surface postPulse/postIterationRecord/postCompletion/setWarmingUp/setRunLabel + new activate(runId). Per-run 5 Hz pulse throttle. resolveWebviewView replays EVERY pane from its buffer (S36) with positional ordering, then posts activate last. setWarmingUp guarded from clobbering a pane that already has data/terminal state. pulse stays plot-only (deliberately does not clear warming). Webview (media/ui/views/inspector.ts): createPanel() instances the former single-run view per runId (no shared globals); a router keys panels by runId, activate toggles the one visible pane, background/late messages only touch their own pane. Pane-hiding uses two-class selectors so it wins over layout.css `.stack` on specificity, not stylesheet order. RunsManager (runs_manager.ts): fans every run's live events into the inspector runId-tagged (routeIter/routePulse ungated); registration replay is state-only so the selected run never double-posts; selectRun adds activate; the single status bar stays selection-gated; completion + promote still fire per-run. Tests: runs_manager + inspector_view_contract updated to the runId-keyed API and fan-out semantics; added per-run isolation, per-run-throttle independence, S36 reopen, activate-last, warming-guard. New happy-dom webview test covers the router itself (per-run isolation, activate toggle, empty-state, plot-only pulse) — closes the coverage gap flagged in adversarial review. All invariants mutation-verified. 121 tests pass; typecheck + build clean. S6 (formulation preview) deferred — no formulation-emit in the frozen contract. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor(1.3): flow RunCompletion whole through completeRun — the #84/#81 seam Jack's #72 merge-seam heads-up: #81 adds `formulation?` to RunCompletion, and completeRun was the third completion path cherry-picking fields positionally (runId/status/fidelity) — once #81 landed, live-completed runs would carry formulation: undefined while replayed runs got it (the exact bug Kate caught on onFinished, reintroduced here). completeRun now takes the WHOLE RunCompletion; both feeders (ingestRunDir's sink verbatim, checkFinished via {runId, runDir, ...readTerminalState()}) funnel the object from the one shared read. An additive field is now a one-place edit (RunCompletion + readTerminalState) and reaches every consumer by construction — consumers cherry-pick at the leaf. Documented as the #84 funnel on both the type and completeRun; the full N-reader consolidation (catalog hydrator etc.) stays #84. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(1.4a): smoke corpus — seconds-scale end-to-end fixtures, Scheduler → executor → run-dir → RunsManager → inspector (#61) Two corpus fixtures (transmon X, cavity displacement — distinct telemetry profiles: 2×8 vs 1×6, 4 vs 3 iters) + test/corpus/fake-julia, a node stand-in the executor spawns exactly like julia (last-argv script, cwd=runDir). It reads each fixture's AMICODE_SMOKE directive and emits the template's telemetry grammar (PULSE_META / ITER / PULSE → run.log via the executor's tail) with a small inter-iter delay so the LIVE tail path is exercised, then writes a schema-conformant result.toml. Zero Julia/Piccolo cost: full chain in ~0.5s. The end-to-end test pins what the unit suites can't — that the pieces AGREE: - Scheduler (#56) lifecycle is strictly serial (B starts only after A's finished event) and satisfies RunsManager's structural seam; - the executor's run-dir writes (run.toml/index/run.log/result.toml/FINISHED) are exactly what the manager's tailer/registry read back (fidelity + iter high-water land in the registry); - telemetry reaches the inspector runId-keyed per run with no cross-tagging (asserted by per-run record dims, completion fidelity, iter records). Runs in the regular vitest suite → already in CI's fast job; #62 promotes it to a named required gate. Mutation-verified: a wrong result.toml fidelity reds the registry + completion assertions. Branch note: contains merge of rchari/56-scheduler (the corpus drives the real Scheduler); diff collapses once #68/#70/#72 land. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(1.4a): review #78 — failure-lane fixture, throwing pumpUntil, wiring-vs-format scope note 1. failing_solve.jl (exit=1): the previously-dead `exit=` directive support now has a fixture — a solve that emits two iterations then dies. Asserts the full failure path end-to-end: executor writes FINISHED{failed}, no result.toml, registry terminal with fidelity undefined but latestIter=2 (pre-crash telemetry tracked), completion fans runId-keyed, promote never fires. 2. pumpUntil now THROWS on timeout with a named condition — a wiring regression fails fast at the offending await instead of an opaque hang. 3. Scope note in the suite header: this guards WIRING (fake ↔ parser), not FORMAT (template ↔ parser) — that boundary is #83's. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> * feat: theme-calculated Harmoniqs yellow — OKLCH-solved brand accent (brand-wide) brand_accent.ts computes the deployed accent from the active theme at webview boot: the canonical #FFF676 ships EXACTLY wherever contrast vs the theme's editor background clears 3:1 (all dark themes); light themes get the closest-to-brand gold by binary-searching lightness with hue + chroma held (gamut-clamped). Two tokens with different jobs: lines (--color-accent, contrast-solved: borders/rings/marks) and fills (--color-accent-fill, always the brand lemon — black text on it ≈ 19:1; a 3:1-darkened gold passes WCAG math but reads muddy under text). --color-on-accent is contrast-picked; yellow is never text. Recomputed live on theme switch. Inspector + catalog-card webviews apply at boot; brand.css statics remain the no-JS fallback. Pill atom gains a dot-less badge variant (dot = process state; badges describe things). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * workbench: run picker + pane-ticker pause + runId-tagged controls (+ Kate's theme accent picked) - amicode.selectRun: QuickPick over the registry (newest first; live/completed/ stopped/failed icons, iter + fidelity + script). Picking pins; "Follow latest" releases the pin via RunsManager.resumeAutoFollow() (jumps to the newest live run). Pre-UX4 utility — unblocks real multi-run testing. - Hidden panes pause their 1 Hz elapsed-strip ticker (Panel.setActive from the router's activate); resumes with a fresh render on re-activation. - Control-row messages carry their pane's runId (post-UX4 correctness; commands still resolve the selected run today). - Cherry-picked Kate's 154650c: OKLCH theme-calculated brand accent — fixes the hardcoded #FFF676 that dies on light themes (audit P0-1, extension side). 408 tests pass; typecheck + build clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * workbench: wire catalog what-next — tune/warm-start stage a concrete chat prompt (clipboard + open chat); promote says Phase-3 honestly Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * workbench: chat theme bridge — iframe boots with ?colorScheme= from the editor theme; live re-theme via onDidChangeActiveColorTheme → two-lane relay (origin-pinned) → app's setColorScheme Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * docs: one-spine mirror breadcrumb — readTerminalState semantics are mirrored in the fork's run-terminal.ts; change both in one change-set Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * workbench: animated H-robot mark in the Run Inspector (replaces the <0||0> text ket; breathe + eye-blink, reduced-motion aware); openRunDir reveals run.toml (bare-dir reveal errors on macOS) with openExternal fallback; inspector auto-open defaults ON Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * workbench: inspector mark = the house silhouette glyph (AmicoSpinner geometry + pulse-opacity language, reduced-motion aware); allow workbench.action.showCommands over the chat bridge (⌘⇧P → VS Code palette) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * workbench: chat auto-opens when the server is ready (amicode.chat.autoOpen, default on) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * workbench: inspector never reveals during the boot index replay — only a run that STARTS while the user works auto-opens it (chat remains the only boot-time surface) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * workbench: boot-quiet test coverage — fresh-run warming asserts the post-boot path; new pin that boot replay never warms/reveals Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * workbench: clipboard bridge — the framed app requests paste over the message bridge; extension answers with vscode.env.clipboard (webviews can't delegate clipboard-read into iframes) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * workbench: open-external bridge — framed app opens https links via vscode.env.openExternal (target=_blank is dead in the webview iframe) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * status bar: stalled-run gate — FINISHED-less run with run.log silent >10min shows 'stalled' (warning), never a perpetual 'running · iter N' from boot replay; mirrors fork isStalled Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * score: never leak the 'issimo' entitlement codename into chat (read as truncated Piccolissimo); dedupe doubled routing paragraph in stage-1 notes Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * stop always terminates: escalation ladder (cooperative STOP → stalled runs get kill+finalize immediately → healthy runs get a 120s grace then an explicit Force-stop offer) - stopPlan: FINISHED → no-op; fresh run.log → cooperative; log cold past the stall threshold (or logless zombie dir) → force - findRunPids: two-key match — cmdline references the run's solve script AND process cwd IS the run dir (sibling runs share the script, never the cwd; no pattern-kills, ever) - forceStop: TERM → 1.5s → KILL survivors → forceFinalize writes the terminal FINISHED (status aborted, atomic rename; run.log breadcrumb) so every contract reader on both spines converges and the UI clears - never a silent kill on a live solver: one long Ipopt iteration can look wedged — the grace path asks before forcing Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(stop): probe lsof at /usr/sbin (macOS) and /usr/bin (Linux) — hardcoded macOS path silently disabled the kill path on Linux (findRunPids proved nothing → force-finalized runs with the solver still alive) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(stop): realpath both sides of the cwd ownership proof — lsof reports physical paths, so a symlinked runs root (/tmp → /private/tmp on macOS) made every pid unprovable and the kill path a no-op Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(stop): forceStop yields to a FINISHED that appeared during the TERM→KILL window — the orchestrator's truthful verdict (failed/143) must not be overwritten with aborted (disk vs registry vs fork endpoints would disagree forever) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(stop): re-prove pid ownership before the SIGKILL sweep — a pid freed by TERM can be reused by an unrelated process inside the 1.5s window Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(stop): JSON-decode script_path from run.toml — values are written JSON-escaped, so escaped chars never matched ps argv and dropped the script-path key from the ownership match Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(stop): pin JSON-decoding of escaped script_path values Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(stop): every stop toast and the force-stop dialog name the run — a nameless dialog 120s later reads as the wrong run being wedged Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(stop): double-stop guard + escalation-timer disposal — a second Stop no longer stacks a second 120s dialog, and the timer dies with the extension instead of firing after deactivate Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(stop): tolerate a deleted run dir — STOP write and finalize are best-effort so a removed dir can't crash the command before the UI entry is cleared Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * perf(runs): 2s TTL cache on liveStatus — boot replay of a long run.log paid one statSync per iter line Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(runs): the poll backstop downgrades the selected run to 'stalled' — routeIter only fires when a line ARRIVES (i.e. not stalled), so a run that wedged mid-watch kept 'running · iter N' forever; downgrade-only so warming/iter flow stays untouched + pin test Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(runs): selectRun consults liveStatus instead of hardcoding 'running' (picking a stalled run stamped a lie nothing would correct); finished runs' display replay no longer flickers stalled/running per line — completion sets the bar once Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: remove empty runs_manager_boot.test.ts accidentally created by a shell append probe two commits ago (vitest fails on suite-less files) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(picker): stalled runs show '$(warning) stalled · iter N' — the picker advertised a wedge as '$(pulse) live', contradicting the status bar Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(bridge): https scheme check is case-insensitive (RFC 3986); clipboard replies gated on panel visibility — a hidden panel rendering LLM-driven content must not sample the OS clipboard in the background Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(scores-e2e): version-agnostic score-marker assertion — hardcoded 'v1' while SCORE.md is v3, red on every creds-bearing machine Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor(runs): one STALL_AFTER_MS (runs_manager imports run_controls's) and the one-spine mirror comment now names the stall threshold + display vocabulary as mirrored semantics Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor(runs): one STALL_AFTER_MS (runs_manager imports run_controls's export); one-spine mirror comment names the stall threshold + display vocabulary as mirrored semantics Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * style: prettier 3.6.2 over the branch's touched files + a minimal .prettierrc (printWidth 120, semi) — the repo had no formatter configured; config formalizes the existing conventions Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * style: prettier over the branch's touched amico-run files (missed in the previous style commit) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * style: prettier repo-wide (new .prettierrc/.prettierignore) — one-time full-repo normalization so the formatter is enforceable from here Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(config): fallback-only model pin — without one, opencode's default resolution gambles on provider ordering and (with Google creds) picked a hanging preview model for every headless/agent turn; anthropic > GA gemini flash, and a user's global model always wins (1.17.3 preserve contract) + tests Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci(vsix-gate): fix by splitting package step — explicit build + fetch:opencode + vsce * test(slow): live-turn extractor is model-agnostic — Gemini opens with the amicode_ask TOOL CALL and no prose, so the ask input (question+options) counts as the turn text; production model pin wired into the e2e servers Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat(bridge): save-file lane — the run-card gallery's PNG export routes through a save dialog (downloads are dead in the framed app); PNG-only, basename-sanitized, size-bounded, relay-allowlisted Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(config): google pin moves to gemini-2.5-flash — the newest flash is capacity-throttled at peak ('model overloaded' → failed turns render as 'model undefined' stubs in chat); the GA flash answered in 1.4s during the same window Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(config): creds-free default model — opencode/deepseek-v4-flash-free (zen free tier, no user quota; answered a tool-bearing turn in ~3s while Gemini kept capacity-throttling); anthropic still wins when creds exist, user global model always wins Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ci: authenticate the private-fork opencode fetch (GH_TOKEN) fast/vsix-gate/boot-smoke fetch the vendored opencode binary from the PRIVATE harmoniqs/opencode release via `gh release download`. The default Actions GITHUB_TOKEN is scoped to this repo only, so gh is unauthenticated for harmoniqs/opencode and the fetch exits 1 — reding every job that vendors opencode (schema-roundtrip is untouched; it needs no binary). Pass a cross-repo token as GH_TOKEN to the three fetch:opencode step defs (boot-smoke's single step covers both matrix legs). Mirrors #80, adapted to this branch's split vsix-gate step. Requires repo secret OPENCODE_FETCH_TOKEN: a fine-grained PAT scoped to harmoniqs/opencode with Contents:Read (validated end-to-end — downloads both assets and the bytes match opencode.lock.json's SHA256 gate). Stays red until the secret exists; green once it's set, no further push needed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(extension): bump opencode.lock to v1.17.3-amicode.2 Pin the vendored opencode binary to the new fork release cut from rchari/amicode-fixes (opencode PR #1): Gemini tool-schema fix, /amicode run-cards + profile endpoints, one-spine run truth, multi-drive pulse fix, save-file bridge. darwin-arm64 + linux-x64 sha256 updated. vsix build verified locally (linux-x64 fetch + sha match + vsce package -> amicode.vsix). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Kate <katebonner277@gmail.com> Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com> Co-authored-by: Jack Champagne <jackchampagne.r@gmail.com> Co-authored-by: Jack Champagne <jack@harmoniqs.co>
Lock gains source=local + ref (merged fork PR #1 head). The fetcher builds from a harmoniqs/opencode clone (sibling ../opencode, AMICODE_OPENCODE_SRC, or --local <path>) after verifying HEAD matches the locked ref and the tree is clean (--any-ref to bypass), then stamps the ACTUAL binary sha256 plus a .source provenance file — replacing the hand-swap workflow that left the stamp lying at the manifest value. No clone present (CI) → falls back to the pinned release download.
Same source (ab0e798) as amicode.2, rebuilt with OPENCODE_CHANNEL=dev so CI and clone-less installs get the amicode UI surfaces by default.
reconciliation: rchari/workbench into problem-workspaces
fetch:opencode: build from local fork clone pinned by git ref
Full package chain + channel-gate check on the vendored binary + packaging-manifest gate, then a GitHub release with amicode.vsix.
ci: tag-triggered VSIX release (v* tags)
opencode.lock: v1.17.3-amicode.4 + ref e9b6951
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.