Repository navigation
Make the policy decision point pluggable (the rules are injectable, the evaluator is not) #350
Description
Activity
The diagnosis holds — I went through all of it while adding a field to
PolicyContextin #526, so
here is what the seam actually looks like from inside, and one thing in the proposal I think is a
harder problem than it reads.The seam is three call sites, and two of them are free
evaluateActionPolicyis called in exactly three places:computer/gateway.ts:532inside async function govern<T>plugins/store.ts:2967inside async callToolcomputer/policy-dry-run.ts:189inside dryRunAgainstHistory, not asyncSo a
PolicyDeciderreturningPromise<PolicyDecision>costs nothing at the two that decide live —
both are already async and already awaiting a provider or a vendor. The injection point is the same
oneoptions.policy()already uses, which is the part of your argument I would keep front and
centre: the data is injectable and the evaluator is a direct import, and there is no reason for that
asymmetry.Worth knowing if you pick this up:
PolicyContextgained aninitiatorfield last week (#526,
{kind, id}— person, deployment, routine or handoff). A plug receives the context, so that is one
more thing it gets for free, and the shape is settled rather than in motion.The hard part is the third call site, and it is not just async plumbing
dryRunAgainstHistoryis how #241 lets somebody test a rule against real history before saving it.
It replays recorded actions through the same evaluator the gateway used, which is the whole reason
the answer means anything.A pluggable evaluator breaks that, and not for a mechanical reason:
- Case 2 in your list — "this bot has already submitted three orders today" — is a decision
about state at a moment. Replaying it against last Tuesday's rows asks today's counter about last
Tuesday's action, and the answer is confidently wrong rather than unavailable. - Case 3 — human in the loop — cannot be replayed at all. You cannot re-ask somebody what they
would have approved a week ago, and a replay that waits on a human is not a replay.
So an evaluator that serves cases 2 and 3 is exactly an evaluator whose answers are not reproducible,
which is the property the dry-run depends on. Making the interface async is easy; deciding what
dryRunAgainstHistorydoes when the deployment has plugged in a non-replayable decider is the
design question.Three ways out, and I have no strong view:
- The interface says so. A decider declares whether it is replayable; the dry-run runs CEL for
the ones that are, and reports "this deployment's decider cannot be replayed" for the ones that
are not. Honest, and keeps the feature working where it can. - The dry-run always runs CEL, explicitly as "what the built-in rules would have said", and
stops claiming to predict the live decider. Simplest, and quietly weakens a feature. - Two seams rather than one — the CEL evaluator stays where it is and a plug is consulted
as well, deny-wins. Keeps replay meaningful, and means a deployment that wants OPA to be the
only authority cannot have that.
Whichever it is, it should be decided before the interface exists rather than after, because it
determines whetherdryRunAgainstHistorytakes a decider at all.One smaller thing
If the point is "a deployment that already owns a policy engine", the fail-closed contract has to
travel with the seam and be stated in the interface, not assumed. Today a thrown CEL expression
denies, a non-boolean result denies, and an absent policy denies —evaluateActionPolicyis careful
about all three. A remote decider adds failure modes the current one cannot have: a timeout, a 500,
a malformed body. If the default for those is anything other than deny, the seam has quietly made
the boundary weaker than the thing it replaced, and that is the kind of change nobody notices until
it matters.- Case 2 in your list — "this bot has already submitted three orders today" — is a decision
What I see today
The policy data is already fully injectable.
options.policy()is a getter(
server/src/computer/gateway.ts:523), backed bycreatePolicyStorewithACTION_POLICY_TOPICfor live reload (server/src/computer/policy-store.ts).The evaluator is not.
evaluateActionPolicyis a direct import from./policy,called synchronously:
So a deployment can change the rules, but not who answers.
Why that's worth a seam
Three cases the current shape can't serve:
A deployment that already owns a policy engine. OPA/Rego, Cedar, an internal
service. Today they have to transcribe their rules into CEL and keep two sources in
sync, which is exactly the setup where the two drift and nobody notices.
Decisions that need state the gateway doesn't hold. "This bot has already
submitted three orders today." "This action is above the amount this actor may
approve alone." "A human granted this, twenty minutes ago, and the grant expires."
CEL over a stateless
PolicyContextcan't express any of those — not a flaw in theexpression language, just a different question.
Decisions that take time. Anything involving a human in the loop is async by
nature, and the current call site is synchronous.
The good news is that most of the work is already done:
PolicyContext(
server/src/computer/policy.ts:44-153) is a clean, serialisable, effect-levelobject —
intent,mcp.effect,file.extension,command, resolvedelement. Thatis a better PDP boundary than most systems that set out to build one.
Concrete shape
Default stays exactly what happens now:
Line 523 is already inside an
asyncmethod that awaitswrite(...)on the nextline, so the
awaitcosts nothing structurally.Two things I'd want written into the contract, not left to the implementer:
matches()and itsonErrorargument (policy.ts:210-243). A remote PDP that ismerely unreachable must not be a way to get an action through, and the timeout should
be required rather than optional.
dry-runstays in the gateway, not in the decider. Otherwise an external decidercan silently disable enforcement for the whole deployment, which is the one thing this
seam must not make possible.
What I would not touch
gateway.ts:469-471). A decider receives theelement the server resolved, never what the caller claimed. The comment at
policy.ts:38-42is the reason the whole thing works and an external decider must notget a chance to weaken it.
gateway.ts:524-535),whoever decided. The trail records the decision, not the decider's opinion of it.
One question before any of this
policy.ts:4-12says this mirrors the policy engine in CopilotKit's enterprise agentgateway, deliberately, so a rule means the same thing in both. Does that engine already
have a pluggable decision point — and if not, is divergence here something you'd rather
avoid? If so, say and I'll drop it; that's a good reason.
Otherwise I'm happy to send the PR: default path byte-for-byte unchanged, the type, the
fail-closed timeout, and one worked example decider.
Disclosure: I build an authorization service for agents, so an external PDP is
obviously useful to me. I've tried to write this as it would stand without that —
push back if it reads otherwise.