Conversation
Signed-off-by: LIghtJUNction <lightjunction.me@gmail.com>
…s around it
A client that wants to say "this command changed these files" has had nothing to
read. The engine records a shell command's writes as a command execution, never
as a `file_change` item, so the only per-call attribution a client could find was
`metadata.mutation`, which only file tools publish. The turn's snapshot delta
covers the turn rather than the call, and deriving it per call meant diffing two
trees the Runtime API did not expose.
The engine already brackets every call that may write — a file tool, a shell
command, a program, a write-capable MCP tool — with a `tool:<call_id>` restore
point before it and a `post-tool:<call_id>` one after, and records both on the
turn. This adds the route that reads that pair, plus the two small pieces it
needs: `SnapshotRepo::has_tree`, so a receipt whose objects have been pruned is
reported as such instead of as a git failure, and `SnapshotRepo::patch_between`,
one path's unified diff between two trees.
GET /v1/threads/{id}/turns/{turn_id}/calls/{tool_call_id}/changes
answers with the change kind, the line counts, the revision the span left and the
patch, per path, all read from the two trees rather than from the working copy.
The answer is therefore what the call did, not what the file holds now. Every
entry carries the span's own revision and the restore point `file-revert` accepts,
so a command's write is revertible like any other.
What the span is not is attribution by cause. It is a window of time: another
writer in the same workspace during the call is inside the difference, and a path
the snapshots do not track (a `.gitignore` hit, `node_modules/`, a binary) cannot
appear at all. A call the engine judged read-only took no receipts, one whose
closing snapshot was lost has half a pair, and an older turn's trees may have
been pruned — each answers `state: "unavailable"` with its own `reason`, which is
a different answer from an empty `files` list and has to be read as "unknown"
rather than "changed nothing".
The tool call id is the model endpoint's own opaque string, echoed back verbatim,
so this compares it rather than parsing it: the charset a *record* id is held to
would refuse `call_01_…|<uuid>` and answer "no such call" for a call the turn
plainly recorded.
Signed-off-by: Ben Gao <bengao168@msn.com>
…dows paths Both platform failures were the same test-side assumption, not a production regression: the suite assumed the raw `TempDir` path is what the config layer reports, which only holds on Linux. The permissions path is resolved through `codewhale_config::normalize_config_file_path`, so the reported path is canonical: `/private/var/...` on macOS and `\\?\C:\...` with the long name on Windows. `crates/config` is untouched by this PR, so this is baseline behaviour that the Linux-only local gate simply cannot see. - `config_policy.rs`: assert the view path against the same resolver the production path uses (`resolve_permissions_path`) instead of the raw fixture path. - `config_policy_baseline.rs`: replace the canonical workspace form before the raw form. Substituting the raw prefix inside the canonical path left a stray `/private` behind on macOS (visible as `"/private<WORKSPACE>/permissions.toml"`) and missed the verbatim path entirely on Windows. Reproduced on Linux with a symlinked `TMPDIR` (raw != canonical): both tests fail identically to the macOS/Windows CI before the change and pass after it.
PR #6832 Windows CI failed only idle_managers_sharing_a_store_do_not_poll_it_continuously: two legitimate startup reloads landed in the zero-reload observation window. The fixed 2.3s startup sleep runs from manager creation, but each worker schedules its 2s metadata retry after its first claim completes. A delayed initial claim can therefore put that retry after the counter reset. Wait for 2.4s of observed load-counter quiescence, with a 10s deadline, before starting the unchanged zero-reload measurement. Apply the same setup to the running-task flush test. Keep task wakeup assertions and production scheduling unchanged. Continuous periodic reloads time out. Validation: - CI-profile nextest, TUI lib, all features: 6 tests run, 6 passed, 14,593 outside the focused selection. - Temporary timing harness extracted the real ClaimSchedule and new wait helper: a 750ms delayed first claim fails the original fixed-sleep assertion, passes with the corrected wait, and an unconditional 2s fallback is rejected at the 10s deadline. - cargo fmt --all -- --check and git diff --check passed. Hosted Windows verification remains pending the follow-up CI run. Signed-off-by: Paulo Aboim Pinto <paulo.aboim.pinto@gmail.com>
…oundaries Replace the PR #6832 idle-test timing workaround with explicit completion receipts on the existing task manager's worker/caller boundary. Each worker owns a bounded inspection mailbox. Every request carries a oneshot reply, sent after the actual queue operation and schedule update. Startup awaits an initial receipt from every worker, with cancellation cleanup, rather than returning merely because workers were spawned. Receipts describe a per-worker decision, not global quiescence or task completion. Claimed tasks execute after their inspection receipt; failed claims return errors and retain their retry/backoff behavior. The three idle/external-queue regressions now use acknowledged cycles and controlled executor start/release messages. Fixtures explicitly supply settled file timestamps, and scheduling tests supply wall-clock inputs. No fixed startup sleep or observation window establishes success in those tests. Existing production metadata-settling and retry policy is retained. Add boundary coverage for all-worker replies, dropped receivers, shutdown, and corrupt-store recovery; adapt the real process-ownership helper to retain control of the worker JoinHandle. Code comments document the completion boundary and why elapsed silence cannot establish it. This repairs pre-existing task-manager code, independently of FEAT-027's command-shape changes. The original regression arrived in 309f329 and its zero-reload contract in 07a6539. Validation: - CI-profile nextest, all features: 79 task-manager/ownership tests passed. - CI-profile nextest: 61 automation/runtime-ownership caller tests passed. - Production TUI library Clippy with the CI lint flags passed. - cargo fmt --all -- --check and git diff --check passed. - Exact extracted scheduling tests: 2 passed; restoring unconditional fallback polling makes the settled-queue test fail immediately. No existing Gherkin scenario owns this internal boundary. The focused suite exercises actual worker messages, persistence, and subprocess ownership/restart behavior. Hosted platform CI remains a separate gate. Signed-off-by: Paulo Aboim Pinto <paulo.aboim.pinto@gmail.com>
Box the plugin provider snapshot so the runtime route remains small enough for async test stacks, and add the contributor credit required by CI. Local checks: - cargo check -p codewhale-tui --lib --locked - cargo fmt --all -- --check - python3 scripts/check-contributor-credit.py - git diff --check
Restore the two large source files after the GitHub API upload truncated their first update. The resulting tree contains only the intended boxed provider snapshot and contributor credit changes.
The image_analyze tool held the original image bytes yet let the vision model guess the image size from the picture, which matters for downstream layout and analysis reasoning. The tool result now carries the image's pixel width and height plus a format label, read with the image crate's header-only image_dimensions probe so no pixel decode is needed; the format label comes from the same extension decision as the mime detection. Animated GIF/WebP report the first frame's size. The probe is best-effort: when the header cannot be parsed — including BMP, whose decoder is not compiled in — the fields are omitted and the tool still succeeds, since the vision request itself does not depend on the metadata. The tool description now tells the model to take image size from that metadata instead of guessing by eye, names the fields as the stored dimensions (camera rotation metadata is not applied), and scopes the promise to images whose container can be sized. Tests: a wiremock-backed end-to-end case asserts real PNG and JPEG dimensions and the format label reach the result payload; an unparsable-bytes case pins the degradation (tool succeeds, no width/height/format keys); a well-formed minimal BMP pins the no-decoder-feature omission boundary. Signed-off-by: asto <asto18089@126.com>
…der LPAC Hosted Windows (LPAC) refuses lstat on the drive root, so Node's JS realpathSync — which walks every ancestor — failed the reviewed-closure checks with EPERM lstat 'C:\' (13 failures on #6815, mostly extension_host::native_mcp and raw_agent_presets). The earlier attempt (4e1905b, reverted) used realpathSync.native but compared its spelling against raw input. New src/dsh/canonical-path.ts is the one comparison rule: canonicalize with realpathSync.native on win32 (GetFinalPathNameByHandle, no ancestor walk) and realpathSync elsewhere, strip \\?\ and \\?\UNC\, compare case-insensitively on win32, and express containment as relative(canonical root, canonical target). A canonical path is never compared to a raw one. resolve-hooks keys closures by canonical root and remembers the raw root; the Node load hook now also refuses symlinked closure modules (under --preserve-symlinks an in-closure symlink pointing outside kept its in-closure URL and loaded — reproduced on HEAD). Evidence (macOS): npm --prefix crates/tui/extension-host test: 406 tests, 399 pass, 0 fail, 7 skipped (new canonical-path test passes); cargo test -p codewhale-tui --lib -- extension_host dsh: 264-265 pass, 2-3 plugins::install::dsh failures that differ per run and pass in isolation (parallel-load flake under investigation). Windows LPAC itself cannot run here; hosted Windows CI is the proof. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Contributor
There was a problem hiding this comment.
Copilot review overview
🔵 Needs a closer look
The security-sensitive Windows LPAC path boundary requires confirmation from the mandated hosted Windows run.
Review effort: Balanced
Findings: None
What changed in this PR
Addresses the Windows LPAC lane by making reviewed extension-host path checks robust to canonical Windows path spellings.
Changes:
- Adds canonical path identity, containment, and symlink checks.
- Updates reviewed hook/composition admission and module loading.
- Adds normalization tests and regenerates extension-host bundles.
| File | Description |
|---|---|
crates/tui/extension-host/test/canonical-path.test.mjs |
Tests Windows and POSIX path normalization. |
crates/tui/extension-host/src/dsh/shell-hooks.ts |
Uses canonical symlink-safe hook checks. |
crates/tui/extension-host/src/dsh/resolve-hooks.ts |
Tracks canonical closure roots and validates loaded modules. |
crates/tui/extension-host/src/dsh/canonical-path.ts |
Implements shared canonical-path utilities. |
crates/tui/extension-host/dist/shell-hooks.mjs |
Regenerates the hook test bundle. |
crates/tui/extension-host/dist/dsh-composition-review.mjs |
Regenerates the composition-review bundle. |
crates/tui/extension-host/dist/codewhale-extension-host.mjs |
Regenerates the main extension-host bundle. |
crates/tui/extension-host/dist/agent-presets.mjs |
Regenerates the agent-presets bundle. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
A foreground command that outlives its wait moves to the background with the output so far and a task id. It is not killed for running long. Jobs lists only real background shells, and the card names the command instead of a Ctrl+B hint. The model is pointed at tool_search and task_shell_wait, because Bash is hidden. A paste or leading drop of an existing image path attaches it. read and read_media can open that exact attached file outside the workspace, and an oversized screenshot is downscaled. codewhale doctor names the environment variable that holds a key, never the value, and a live probe decides the verdict. Validation: direct rerun of the lane-1 test binary, 13 passed and 0 failed (lowercase bash description and detach, foreground move-to-background, origin identity, doctor_verdict_tests). A broader filter of image_attach, shell, history, jobs and doctor tests was green before a 60-second work-surface test was still running when the watch ended. Conformance goldens re-recorded for the changed tool descriptions. cargo fmt clean on the lane files. clippy -D warnings not run for this slice.
A developer home keeps one built-in computer-use record and snapshot per installed build. /plugin doctor reports what is stale, and --fix retires those records and snapshots, then compacts state.json. The previous file is kept as state.json.pre-gc. It does not delete a user plugin that still has a bundle directory, a snapshot a running process names, or anything inside the grace window. Until this build has its own record, the review it would carry forward is kept. Validation: cargo test -p codewhale-tui --lib -- doctor_reports_read_only_then_fix_retires_stale_records_with_a_backup -- --exact: 1 passed, 0 failed. An earlier gc filter reported 16 passed and this one failure, which was the test looking for an unescaped hyphen in escaped output.
An HTTP 400, 402, 405, 409, 413 or 422, a spent balance, a context-budget stop, and a turn's own step or wall-clock ceiling now carry an input or budget label. A bare ERROR is an unreadable provider error, not a warning. An in-200 error frame that arrives before any content is retried, and its recoverable flag tells the truth. Refs #6843 and #6795. The wider input category must not make compaction drop history. A max_tokens or temperature rejection names a token or a maximum without being a length overflow, so the drop-oldest ladder now requires real overflow wording. Validation: an earlier cargo test of error_taxonomy, error_frame, placeholder_error, exhausted_transient, midstream_error, classify_error, session_diagnostics and termination reported 47 passed and 0 failed. Not re-run after the later compaction wording guard.
The drop-path sniff and the attached-image admission each stat and open or canonicalize. The read path already runs that admission inside spawn_blocking. The paste and submit checks run on the synchronous UI thread, where a blocking stat cannot park a Tokio worker. The budget records the sites instead of pretending they are wrapped. CI counted 6 against a budget of 5 on PR 6846.
… theft restore_compaction_checkpoint anchors at the first provenance-stamped carrier and, on the wire predicate, never treats user text as a checkpoint. The existing duplicate-carrier test covers a pasted summary appearing AFTER the real carrier; the mirror order - pasted full summary BEFORE the carrier - is the shape that a content-based anchor gets wrong, because the pasted header matches first and steals the insertion position while the real carrier is replaced elsewhere or dropped. Port that mirror-order regression: with a pasted summary ahead of the stamped carrier, restore must keep the pasted turn verbatim as user content, anchor at the carrier's index, and replace the carrier there with the saved summary. The carrier's text differs from the saved summary so the assertions cannot pass by leaving the history untouched, and the pasted turn carries no provenance block so the assertions fail under any content-based anchoring. Signed-off-by: asto <asto18089@126.com>
/plugin doctor hashed plugin directories with path.canonicalize() on the caller. That caller is the TUI thread, which is a Tokio worker, and the blocking-call ratchet failed the four new sites. Grok Build keeps a filesystem walk synchronous and moves the resolving once, in one spawn_blocking, instead of wrapping each lookup. The doctor now does that: one task resolves the home, lists that home, and resolves what it found. A path the task did not resolve is kept, not retired. python3 scripts/check-blocking-calls-budget.py: exit 0, 534 sites across 165 files, within budget. cargo test -p codewhale-tui --lib plugins::registry::gc::tests: 16 passed, 0 failed, 1 ignored. commands::groups::plugins::tests::doctor_reports_read_only_then_fix_retires_stale_records_with_a_backup: 1 passed, 0 failed.
Version drift failed on #6846 because crates/tui/CHANGELOG.md is a generated slice of the root file, and the root Unreleased section was empty while the slice already named /plugin doctor, #6843 and #6795. The notes now live in the root file. ./scripts/sync-changelog.sh --check exits 0 and the generated slice is unchanged.
A command still running after its foreground wait moves to the background and is not killed. The completion hooks were still expecting the old timed_out failure, which is what Test (ubuntu-latest) failed on 804c631: left "call-timeout unset running true", right "call-timeout unset timed_out false". The after-hook now expects running, and on_error no longer fires for that case because the result is a success. Nonzero exits stay failures. The shell description that says this grew every tool surface by 564 schema bytes and 141 estimated tokens. No tool entered or left a surface. The runtime-contract ceiling records that growth. cargo test -p codewhale-tui --lib -- runtime_shell_completion_delivers_exit_code_and_status_to_hooks bash_completion_hooks_get_exit_code_and_status_for_failures: 2 passed, 0 failed. python3 scripts/check-runtime-contract-budget.py: PASS, all 55 metrics exactly at budget.
Preserve the wave wording and its own exact measured contract: provider-free runtime budget passes at 7903b96; main repair is independently measured from hosted Linux. This merge adds only its history note, with no numeric ceiling or Rust source change relative to wave. npm gate remains 1285 passed, 0 failed, 7 skipped; check:web passed. Budget measurement 4/4 exact metric probes, checker/measurement tests 25/25. No new native or hosted qualification claimed. Signed-off-by: CodeWhale Bot <bot@codewhale.net>
Correct the mistaken pass claim in a53b58b: the exact provider-free wave receipt failed four stale full-catalog ceilings. Measured Act/Operate full is 85359 bytes / 21340 estimated tokens (+644 B / +161). Merged documentation adds 714 B; the common full-surface factual rewrites trim 70 B (Plan full 54336 -> 54266). Raise only those four fields; other numeric ceilings and all identities pass. The same saved receipt now passes the checker. No second compile or full suite rerun needed for data-only correction. Evidence: 4/4 exact runtime metric probes; 25/25 budget checker/measurement unit tests. Unchanged JS gate inputs: npm test 1285 passed, 0 failed, 7 skipped; check:web passed. Actual macOS receipt is in CW/artifacts/codex-engine-takeover-20261005/wave-runtime-contract.json. Linux final-head check remains required. Signed-off-by: CodeWhale Bot <bot@codewhale.net>
…atch The canonical CLI selected Stdio/Socket/LegacyHttp in typed startup options, but Serve discarded it and always constructed control_frontend=None. app-server --stdio therefore printed an HTTP endpoint and never answered stdin. Pass the captured selection through to the Runtime API owner. Bound and isolate the release stdio probe: canonical short temporary home for no-follow credential directories and macOS Unix socket limits, no ambient provider credentials, no telemetry, 20-second deadline, all six successful JSON-RPC replies and clean process shutdown required. Evidence: old 5b3f4b3 release binary fails the new probe at its deadline; fixed release build passes 6 checks, 0 failed, and shutdown exits 0. Smoke script tests: 22 passed, 0 failed (deadline and credential isolation). Release build --locked passed (11m35s). npm test: 1285 passed, 0 failed, 7 skipped; check:web passed. rustfmt, shell syntax, git diff --check pass. Blocking budget 534 sites/165 files and dead-code budget 253 <=258 pass. Real PTY startup and Ctrl+D exit passed without a provider prompt. Hosted Linux/Windows qualification remains pending GitHub runner allocation. Signed-off-by: Hunter Bown <hunter@hmbown.com>
…auth /v1/thread-history/operations/lookup, /v1/thread-history/operations/recover, /v1/thread-history/mutate and /v1/thread-history/import were mounted on the public router, outside require_runtime_token + require_workspace_scope, so any loopback client could mutate, import, or recover thread history without the bearer token every other mutating route requires. Move all four into api_routes behind the existing middleware; /v1/runtime/info stays public by design (loopback clients discover the listener there — documented in place). Regression test asserts anonymous and wrong-token POSTs to all four routes return 401 and a correct token does not. Evidence: cargo test -p codewhale-tui --lib --locked thread_history -> 5 passed; 0 failed, including runtime_api::tests::thread_history_operation_routes_require_auth.
… disclosed The workspace-write posture on Linux silently ran unrestricted: prefer_bwrap defaulted to false, so select_sandbox resolved SandboxType::None and the confined command became an unconfined one. Reproduced as an outside-workspace write succeeding under a workspace-write policy. Three parts, mirroring the extension host's proven pattern: - Config::prefers_bwrap() resolves the key as on-by-default; every consumer (engine config, exec agent, TUI init/frame, runtime threads, doctor/setup status) migrates to the accessor. prefer_bwrap = false remains a working opt-out and danger-full-access is untouched. - bwrap::is_available() now runs a real wrapped /bin/true once (built by build_bwrap_command itself, cached in a OnceLock). The exec bit alone lied on hosts that restrict user namespaces (Ubuntu 24.04 apparmor_restrict_unprivileged_userns), where every wrapped command would fail. - codewhale exec prints a stderr warning when the resolved policy wants a sandbox but no enforcing wrapper exists, so policy-only never goes silent for the human. The model already saw the same posture via turn_meta. Real Linux proof, not a marker check: new test bwrap_workspace_write_blocks_outside_write_allows_inside runs the same outside-workspace write unwrapped (lands on the host) and wrapped (EROFS, file absent), plus a normal inside-workspace write that succeeds under the wrapper. Evidence: cargo test -p codewhale-tui --lib --locked -- bwrap prefers_bwrap linux_ parity_manager -> 26 passed; 0 failed. Docs aligned across SANDBOX.md (en/zh), CONFIGURATION.md (en/zh), config.example.toml (also drops the stale Landlock comment), the FAQ page and dictionaries, features.toml, and the public-surface facts both sides of the drift gate.
# Conflicts: # crates/tui/src/snapshot/repo.rs
clap ate values like -y as flags. allow_hyphen_values keeps `mcp add srv --command npx --arg -y` intact. Evidence: cargo test -p codewhale-tui --lib --locked mcp_add_arg -> 1 passed (mcp_add_arg_tests::mcp_add_arg_accepts_hyphen_values).
On account sign-in success, when the local provider is still the default, select provider codewhale + model auto and save, so chat works immediately. An explicitly configured provider is left alone with a keeping-your-route notice. Evidence: cargo test -p codewhale-cli --lib --locked cloud::tests::login_ -> 2 passed (selects_managed_route_when_default, keeps_explicitly_configured_route).
list_call_changes resolves the side repo on a blocking-pool thread. Under main's sealed test home that thread is a foreign reader: it resolves the isolated snapshot root instead of the test's store and every span reads as pruned. Mint the scope ticket on the handler and join it on the worker, the same pattern file-revert already uses. Evidence: call_change_route_reads_one_calls_workspace_span ok; patch_between/has_tree/run_git_drains neighbors 4/0.
The #6827 Windows gate refused -Id $var even when the same command assigned the variable from a source the gate accepts inline (literal PID, port-owner lookup, Start-Process -PassThru), blocking the start-then-stop-your-own-server pattern the refusal recommends (#6871). A variable assigned exactly once from one of those shapes is now usable at -Id/-Id:/taskkill /PID ($var.Id for -PassThru); any second assignment, compound op, ++/--, pipeline, or wrong accessor keeps the hold. Proofs stay off for interact stdin, which runs in a persistent session earlier payloads may have primed. Evidence: cargo test -p codewhale-tui --lib --locked -- <4 gate tests> -> 4 passed (held-table + allow-table + stdin-scope).
An outstanding request_user_input now owns its wait beyond the UI tool watchdog, including nested execute_tools and temporarily missing tool cells. Answering the request retires the exemption and preserves stalled-tool recovery. Fixes #6872. The Linux bwrap proof explicitly excludes temporary writable roots: hermetic CI HOME lives under /tmp, which the default policy intentionally permits. Production sandbox policy is unchanged. Restore rustfmt for the MCP hyphen argument regression and keep English/Chinese sandbox code blocks identical with their GT catalogs. Contributors can build stamped current main source from the canonical codewhale-hq repository; correct the minimum Rust version to 1.89 and point to focused local tests. Validation: hermetic nextest turn_liveness + mcp_add_arg: 12 passed, 0 failed (15166 unrelated filtered). npm test: 1285 passed, 0 failed, 7 skipped (wrapper 101, SDK 19, extension host 399/406, web 766). npm run check:web passed. cargo fmt --all -- --check and git diff --check passed. Linux enforcement awaits hosted Linux CI; no release or deployment claim.
An unanswered human question stays live independently of the TUI hung-tool timer. Retire only the matching question on tool completion, turn end, cancellation or disconnect, including hidden cards. Give a successfully accepted answer fresh turn activity so the completed human wait cannot immediately trigger the watchdog. Engine submit/cancel acknowledges the current live request before its configured deadline or cancellation; a mailbox enqueue alone no longer reports success. Bound host delivery to five seconds after the decision, reject abandoned replies, retain the Runtime settlement claim until the verdict, and restore the same durable question on failed delivery. Terminal completion prevents an in-flight rejected answer from resurrecting a stale waiter. Durable intent remains recorded before Engine consumes the answer. Validation: hermetic nextest focused lifecycle/acknowledgement/API/liveness suite 90 passed, 0 failed (15100 filtered). npm test 1286 passed, 0 failed, 7 skipped: wrapper 101, SDK 19, extension host 399/406, web 767. npm run check:web passed. Formatting and diff checks passed. Hosted cross-platform CI and a real DeepSeek terminal session remain separate proof. Refs #6872 Signed-off-by: CodeWhale Bot <bot@codewhale.net>
…er images Point Rust and TypeScript default GitHub report routing to codewhale-hq/Codewhale, regenerate the extension host and its reviewed harness digest, and update parity fixtures together. Container guides use the verified public ghcr.io/hmbown image namespace while distinguishing published v0.10.0 from a current-source build; the new org image has not been published yet. Validation: hermetic focused GitHub/harness Rust suite 49 passed, 0 failed; GitHub adapter Node tests 10 passed, 0 failed; npm test 1286 passed, 0 failed, 7 skipped and npm run check:web passed. Public image manifests for latest and v0.10.0 returned 200 with the same digest; new-org publication remains separate. English/Chinese examples checked together. git diff --check passed. Signed-off-by: CodeWhale Bot <bot@codewhale.net>
Use one authoritative typed Cloudflare config, preserving the existing Worker name, KV identities, Durable Object export, bindings, domains and cron. Generate the ignored OpenNext adapter config from it. Pin cf beta.12 and its official internal build delegate; build and populate OpenNext cache before cf deploy --prebuilt. Fail closed if deployment SHA or storage identities are missing. No production deployment, namespace creation or secret change was performed. Exclude generated cf/preview bundles from the source locale ceiling and fix two vulnerable transitive parser versions. The recovered terminal gallery remains explicitly a 0.10.1 pre-release preview with 0.10.0 the published release. Validation: npm test 1286 passed, 0 failed, 7 skipped (wrapper 101, SDK 19, extension host 399/406, web 767); npm run check:web passed. Focused deployment preflight 13/13 passed, TypeScript passed, OpenNext plus cf bundle build passed, cf deploy --dry-run passed, KV identity guard passed. Real local Worker /api/facts returned 200; browser interaction selected and rendered all six terminal views. Narrow viewport had no horizontal page overflow. No deploy/publication claim. Signed-off-by: CodeWhale Bot <bot@codewhale.net>
feat(runtime-api): read what one tool call changed, from the snapshots around it
fix(search): honor the configured locale in Bing and DuckDuckGo scrapes
feat(vision): report real image dimensions in vision results
feat(automation): archive terminal runs on delete
feat(runtime-api): serve a skill's body so a client can activate it
…eft-test test(compaction): pin restore anchoring against pasted-summary anchor theft
Reconcile the complete contributions from #6832 (@aboimpinto), #6867 (@hodeswildsmith455-boop), and #6805 (@LIghtJUNction) with the current Engine, provider identities, reviewed-plugin policy, and task lifecycle. Original contributor histories are recorded by subsequent resolved merges. Fix the already integrated contributor cases: snapshot corruption is an explicit unavailable/error result (#6817), malformed locales cannot select an incidental script (#6860), image metadata uses the exact uploaded bytes and respects available decoders (#6858), automation deletion waits for actual scheduler reconciliation (#6864), and blocking trust/skill reads stay off the async executor (#6869). Preserve #6857's compaction regression. Extend #6872's human-wait lifecycle guard to approval/elevation cards; retire only the matching ended parent request and refresh activity only after a delivered decision. Enforce configured finite approval deadlines in the Engine, including deadline/cancellation races and durable receipts. Serialize Native Windows ACL admission/retirement across Core processes with a logon-scoped kernel mutex, preserving exact SID/object validation. Add an actual child-process lock test; serialize DSH host tests in the existing extension-host lane. Windows execution proof remains hosted CI. Reconcile the vendored computer-use plugin with canonical main a656f67455fc while preserving Core's 0.12.1 embedding contract. Canonical b47/a656 tree passed Ubuntu/macOS/Windows source and package gates, including 28/28 Windows-focused tests and the controlled desktop fixture; this is separate from the new Engine head's CI verdict. Partial adaptation of the discovery-cache priority portion from PR #6393 by @AdityaVG13 (original ac33dd4). Preserve the best match under count and byte limits without importing the unfinished echolocation/fork design; the broader draft remains open. Validation: - npm test: 1286 passed, 0 failed, 7 skipped; web 767/767. - npm run check:web: lint, typecheck and production build passed. - Affected Rust selection: 801/803 initially passed; the two fixture/lifecycle expectation failures were corrected and each passed a focused rerun. - Additional focused Rust: approval 28/28, discovery cache 11/11, OrcaRouter synthetic catalog 3/3, and 33/33 lifecycle/API/routing checks. - Final CI-repair selection: 37/39 initially passed; the BMP feature-proxy and feature-registry summary failures were corrected; both corrected tests passed (2/2, 0 failures). - Qualified Clippy: six packages, all targets/all features, passed with the CI style allowances. Portable no-default-feature check passed; portable policy verifier 6/6 passed; formatting and diff checks passed. - Runtime contract: 55 measured metrics passed after explicit remeasurement; all 21 structural identities were unchanged. Twelve byte/token-estimate budgets account for bounded child-wait disclosure and contributor locale descriptions; no runtime field/prompt was removed to lower the budget. - Persistence budget was not qualified locally: the checker requires a clean tree and this shared checkout retains an unrelated operator file. Clean hosted CI, fresh stamped build and real DeepSeek TUI acceptance remain required before the integration PR can merge to main. Refs #6872, #6843, #6795. Co-authored-by: Paulo Aboim Pinto <paulo.aboim.pinto@gmail.com> Co-authored-by: hodeswildsmith455-boop <hodeswildsmith455-boop@users.noreply.github.com> Co-authored-by: LIghtJUNction <lightjunction.me@gmail.com> Co-authored-by: AdityaVG13 <adityavgcode@gmail.com> Signed-off-by: CodeWhale Bot <bot@codewhale.net>
Record the original contributor history after complete manual reconciliation in b0fb5f8. The tested source tree already includes this PR's intent with current Engine/API/policy conflicts resolved; this merge records the resolution without replacing that verified source tree. The original contribution remains attributed to its original author(s). Signed-off-by: CodeWhale Bot <bot@codewhale.net>
Record the original contributor history after complete manual reconciliation in b0fb5f8. The tested source tree already includes this PR's intent with current Engine/API/policy conflicts resolved; this merge records the resolution without replacing that verified source tree. The original contribution remains attributed to its original author(s). Signed-off-by: CodeWhale Bot <bot@codewhale.net>
Record the original contributor history after complete manual reconciliation in b0fb5f8. The tested source tree already includes this PR's intent with current Engine/API/policy conflicts resolved; this merge records the resolution without replacing that verified source tree. The original contribution remains attributed to its original author(s). Signed-off-by: CodeWhale Bot <bot@codewhale.net>
added 2 commits
October 6, 2026 00:36
Record the actual embedded manifest version and canonical source revision in the current release notes; regenerate the packaged TUI changelog. Validation: bundled-plugin claim check passed; version state and 38 feature references passed. npm test: 1286 passed, 0 failed, 7 skipped (web 767/767). npm run check:web: lint, typecheck and production build passed. The previous hosted Version drift failure was a missing version receipt, not permission to waive the gate. Fresh exact-head CI remains required. Signed-off-by: CodeWhale Bot <bot@codewhale.net>
Request the canonical codewhale-hq repository and match the canonical URLs GitHub now returns. Keep exact tag, asset, digest, receipt and notarization checks; old-owner and lookalike metadata remain refused. Verified against the real published v0.12.0 GitHub release and release.json: the previous checker returned pending; this checker returns ready with matching ZIP and DMG digests and receipt sizes. No release was published. Validation: npm test 1289 passed, 0 failed, 7 skipped (web 770/770); npm run check:web lint, types and production build passed; diff check passed. Three new canonical-URL and noncanonical-owner cases passed. The installed 903b26e CLI has exactly the same Rust workspace and embedded assets as this commit; only these two website files differ. Signed-off-by: CodeWhale Bot <bot@codewhale.net>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Codewhale 0.10.1 integrates the contributor release queue and fixes human waits that could be mistaken for hung tools. A question can remain unanswered with a zero timeout; a later answer must reach the original live turn. Ended or cancelled requests retire their own cards, and finite approval deadlines are enforced by the Engine.
Contributor work retained as original PR histories:
Additional release work:
cf; parser patches and exact-main deployment checks are included. Production has not been deployed by this work.Validation before the final source push:
npm test: 1286 passed, 0 failed, 7 skipped (web 767/767).npm run check:web: lint, types and production build passed.Acceptance still required before main merge: this exact head's Linux/macOS/Windows, lint, safety and web CI; a fresh stamped build; real DeepSeek terminal acceptance including an unanswered question held beyond ten minutes. Signing, release publication and production website deployment remain separate gates.
Refs #6872, #6843, #6795.