feat(supervisor-middleware): add network egress middleware#2027
Conversation
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
595191e to
97b750f
Compare
358906a to
1fbcdbc
Compare
|
🌿 Preview your docs: https://nvidia-preview-pr-2027.docs.buildwithfern.com/openshell |
This comment was marked as outdated.
This comment was marked as outdated.
This comment was marked as outdated.
This comment was marked as outdated.
2 similar comments
This comment was marked as outdated.
This comment was marked as outdated.
This comment was marked as outdated.
This comment was marked as outdated.
|
/ok to test c4b0dcf |
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
…are outages An unreachable operator-registered middleware service previously aborted sandbox startup via a hard error in load_policy, contradicting the per-request on_error contract and the resilient live-reload path. Retry the initial connect and, on failure, degrade to the built-in registry so matched requests are governed by each config's on_error (deny for fail_closed, allow for fail_open) instead of blocking the whole sandbox. The policy poll loop now reconciles the registry on every poll while an install is pending, so a recovered service is adopted without waiting for a config change; a failed reconcile also no longer blocks unrelated policy updates. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
…limit A chain entry whose binding did not resolve reported a zero body limit, which dragged the whole chain's buffer cap to zero and spuriously failed body-bearing requests over capacity even when a resolved middleware could have processed them. Exclude unresolved entries from the limit via a new DescribedChainEntry::is_resolved(); when no entry resolves, skip buffering and apply each entry's on_error directly. Also fix two parallel-test flakes found while validating the change: - Build middleware OCSF events into a Vec and assert on it directly instead of capturing through the global tracing pipeline, whose callsite-interest cache is process-global and raced under parallel runs. - Accumulate the websocket deny response until the reason marker arrives rather than assuming a single read returns the full body. Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
c4b0dcf to
2b7cf4e
Compare
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
|
Addressed the latest findings in
Validation:
|
BlockedThanks @pimlock. I checked your July 16 update on head Gator is blocked by merge conflicts against Head SHA: GitHub currently reports Next action: @pimlock needs to rebase or merge |
Signed-off-by: Piotr Mlocek <pmlocek@nvidia.com>
|
Merged current The single conflict was in Validation: |
Re-check After Author UpdateThanks @pimlock. I re-evaluated latest head What I checked: the current merge commit, the combined Disposition: resolved. The prior merge-conflict blocker is resolved ( Remaining items:
Next state: |
Monitoring CompleteMonitoring is complete because this PR has merged. Final status: the last active gator state was I removed the active |
Resolve conflicts with the corporate HTTP proxy work (NVIDIA#2245), the CLI commands/common extraction (NVIDIA#2359), the supervisor middleware runtime (NVIDIA#2027), and the shared L7 endpoint validation refactor (NVIDIA#2389): - openshell-core/src/net.rs: keep both import sets — main's `ipnet` CIDR types alongside `SocketAddr`/`TcpStream` for the TCP_NODELAY helpers. - openshell-cli/src/run.rs: main moved the `progress` constants into the test module; keep only the `net` import at the top level. - openshell-server/src/lib.rs: union of both import lines. - supervisor-network/src/proxy.rs: main replaced the CONNECT tunnel's direct dial with proxy-aware `dial_upstream`. Take main's call and push TCP_NODELAY down into `dial_upstream`'s direct paths plus `upstream_proxy::connect_via_inner`, so tunnel hops keep it whether or not they chain through a corporate proxy. The plain-HTTP forward path keeps main's direct-dial comment with the NODELAY-setting connect. Signed-off-by: Jim Meyer <jim@meyer4hire.com>
Summary
Implements the first usable RFC 0009 supervisor middleware slice: proto-backed, host-selected HTTP egress middleware for
HttpRequest/pre_credentials, with both in-process built-ins and statically registered operator-run gRPC services.The implementation covers RFC 0009 Phase 1 and adds basic external-service support from Phase 2. It establishes the contract, policy plumbing, ordered chain execution, built-in secret redaction, static gateway registration, relay integration, validation before policy persistence, body limits, audit events, and user-facing configuration and operations documentation.
Related Issue
Closes #2010
Part of #1733
Design/RFC: #1738
Changes
openshell/secretsredactor and statically registered operator-run gRPC services.network_middlewarespolicy configuration and validation, independent of the network policy rule that admits a request.Testing
mise run pre-commitpassesChecklist