Repository navigation
Reject durable replay and HITL review with a SandboxToolset, and fail the task when a sandbox cannot be provisioned - #73529
Merged
Conversation
…olset, and fail the task when a sandbox cannot be provisioned A sandbox is destroyed when the agent run ends, so two operator features that assume otherwise produced wrong answers instead of errors. durable=True replayed cached tool results such as 'wrote the file' against a sandbox that no longer existed, then ran the first cache miss against a fresh empty one. enable_hitl_review=True regenerated after feedback in a second run whose message history described files the first run's sandbox had held. AgentOperator now refuses both at construction, looking inside prefixed, filtered and combined toolsets and inside Toolset capabilities, and names the two ways out. A backend raising a recoverable SandboxError from create() used to escape the toolset's error mapping and fail the task by accident. It now fails it on purpose: the model has no input into provisioning, so nothing it retries can help, and Airflow's own retry is the right one. The create() contract says so. Also covers the liveness probe branch where a non-Modal error keeps the handle.
vatsrahul1001
approved these changes
Sep 22, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A
SandboxToolsetprovisions its sandbox on the model's first tool call and destroys it when the agent run ends, so nothing in it survives into a task retry or a second run. TwoAgentOperatorfeatures assume it does, and until now each of them produced a wrong answer rather than an error:durable=Truereplays cached tool results on a retry without calling the backend. A replayedwrite_filereported success while no sandbox existed, and the first call that missed the cache ran against a fresh, empty one. The model was handed a filesystem that did not match what it had just been told, and nothing raised.enable_hitl_review=Trueregenerates after reviewer feedback by starting a second agent run. That run got an empty sandbox while its message history still described the files the first run had written.The docs called both combinations unsupported; this makes the operator refuse them at construction, the same way it already refuses
durablewithcode_modeand withenable_hitl_review. The error names the two ways out: drop the flag, or move the sandbox work into its own task.The check looks inside compositions, not only at the top level of
toolsets=..prefixed(),.filtered()and.prepared()each wrap the original toolset, several toolsets passed together become aCombinedToolset, and tools can also reach the agent through aToolsetcapability inagent_params. Running two sandboxes on one agent, which the docs recommend, goes through.prefixed(), so a shallowisinstancewould have missed exactly the documented shape. The walker is a small helper inutils/toolsets.pywith its own tests.A failed provisioning now fails the task on purpose.
call_toolawaited_ensure_sandbox()outside thetrythat maps a recoverableSandboxErrorto aModelRetry, so a backend raising the recoverable class fromcreate()had it propagate untouched and fail the task by accident, while the contract said the model could work around it. The alternative was to move the await inside thetryand let the model retry provisioning, but the model has no input intocreate(): it takes only the spec, which is fixed in the Dag file, so no retry the model makes can turn a bad image tag or a rejected credential into a working sandbox, and letting it try would spend its retry budget on a fact it cannot see. The toolset now re-raises a recoverable create error asSandboxTerminalErrorwith the cause chained, so Airflow's own task retry attempts the provisioning again, and thecreate()docstring states the rule. The Modal backend already behaved this way deliberately;sbxraised the terminal class from most create paths, so the gap was latent there.Also adds the test a reviewer asked for on #72910: the Modal liveness probe keeps the handle when the probe itself fails with something other than a Modal error, since that says nothing about the sandbox and the command it followed had already produced its output.
Behaviour change. A Dag that combined
durable=Trueorenable_hitl_review=Truewith aSandboxToolsetused to parse and run, wrongly; it now fails at parse time with the message above. A toolset resolved per run from a callable, such as aToolsetcapability holding a factory, cannot be inspected when the operator is built, so it is the one composition the check does not see. The docs say so.Two imports move from function bodies to the top of
operators/agent.py. pydantic-ai is already imported at parse time through the hook, and the sandbox toolset does not import the operator, so there is no cycle. The Modal backend already wrapped a recoverable create error into the terminal class itself, andsbxraises only the terminal class fromcreate(), so neither backend changes behaviour here; the fix is for the contract and for backends written against it.{pr_number}.significant.rst, in airflow-core/newsfragments. You can add this file in a follow-up commit after the PR is created so you know the PR number.