Repository navigation
Add human approval for agent tool calls - #73586
Merged
Merged
Conversation
phanikumv
approved these changes
Sep 23, 2026
kaxil
marked this pull request as draft
September 23, 2026 10:40
kaxil
force-pushed
the
common-ai-tool-approval
branch
from
September 23, 2026 10:46
2b252fc to
dcf5be4
Compare
kaxil
force-pushed
the
common-ai-tool-approval
branch
from
September 23, 2026 19:14
dcf5be4 to
99fa0b8
Compare
Tools marked with pydantic-ai's approval API (toolset.approval_required() or Tool(requires_approval=True)) now pause AgentOperator / @task.agent in the awaiting_input state and ask on the Required Actions page. Approve runs the call and the agent carries on; Reject tells the agent why and it carries on without the call. A task instance asks once per Dag run. Needs Airflow 3.3+.
kaxil
force-pushed
the
common-ai-tool-approval
branch
from
September 28, 2026 00:20
99fa0b8 to
2e532c2
Compare
kaxil
marked this pull request as ready for review
September 28, 2026 11:48
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
An agent that can look things up freely should not refund an order, send an email or write to a production table without a person seeing the call first. Mark those tools with pydantic-ai's own approval API, and
AgentOperator/@task.agentpause the task in front of the call and ask on the Required Actions page, with the tool name and its arguments:Approve runs the call and the agent carries on. Reject does not fail the task: the agent gets the reviewer's reason as the tool result and carries on without the call. While it waits, the task is in
awaiting_inputand holds no worker slot. Needs Airflow 3.3+.Design rationale
Why not
require_approvalorenable_hitl_review? Both review an output:require_approvalgates anLLMOperatorresult after the run, andenable_hitl_reviewreviews an agent's final answer while polling in the worker. Neither can stop a side effect before it happens, which is what a tool call needs.Marking is pydantic-ai's; Airflow supplies the backend.
toolset.approval_required()andTool(..., requires_approval=True)already exist and make pydantic-ai end the run withDeferredToolRequests. The operator addsDeferredToolRequeststo the output type only while building the agent. pydantic-ai strips it from the schema the model sees, andself.output_type(part of the serialized Dag) is unchanged. The pause is the sameTaskAwaitingInputpathLLMApprovalMixinuses on 3.3+, and the resume continues the run from its transcript withDeferredToolResults.The transcript goes to the task state store, not the continuation kwargs. It holds tool results such as query rows, which do not belong on the task instance row. The continuation carries the transcript's hash, the pending call ids, the usage so far, and the rendered toolset ids. The resume fails if the transcript changed, or if a templated connection id now renders to a different connection than the reviewer saw.
One approval per task instance per Dag run. Core keeps a single approval request per task instance, across retries and clears, and a repeat request keeps the first one's subject and body (
execution_api/routes/hitl.py). A second request would show the reviewer the earlier call while Approve applied to the new one. So the operator records a marker in the task state store, and a second request fails with a non-retryableToolApprovalAlreadyRequestedError. Allowing several approvals needs core to overwrite the request content on a repeat upsert; that is a separate change.No approve-on-timeout.
on_tool_approval_timeoutis"fail"(default) or"deny", and"deny"tells the agent that nobody answered, not that a person refused.usage_limitscovers both sides of the pause, so acost_limitis not reset by it.Screenshots
On a real scheduler and API server with a
testmodel: two runs waiting on the same tool, answered through the UI.Approve with the reason left empty: the refund runs and the agent finishes. The resumed run's
run_idcarries the-resumedsuffix.Reject with a reason: the agent receives it as the tool result, and the task still succeeds.
Gotchas
durable=True,enable_hitl_review=True,code_mode=Trueor aSandboxToolset, each of which assumes the run finishes in one go. There, a marked tool fails the task as it does today, and addingDeferredToolRequeststooutput_typeby hand raisesUnsupportedToolDeferralError.CallDeferred) are not supported and fail without retrying.run_idalready present in the message history, so the resumed run uses<task-instance id>-resumed. Therun_idXCom holds that id;usagecovers both sides of the pause.LoggingToolsetnow logs a call waiting for approval at INFO rather than as an ERROR "Tool ... failed".LLMOperator's assigned-users validation moved into a helper shared with the newtool_approval_assigned_users; its behaviour and messages are unchanged.{pr_number}.significant.rst, in airflow-core/newsfragments. You can add this file in a follow-up commit after the PR is created so you know the PR number.