Repository navigation
Add a Modal backend for the sandbox toolset - #72910
Merged
Merged
Conversation
SandboxToolset has shipped with one backend, sbx, which needs the sbx binary on the worker host and KVM on Linux. That rules out an unprivileged container, so every Kubernetes deployment has been unable to use the toolset at all. ModalSandboxBackend runs each sandbox in Modal instead. Nothing is installed on the worker, model-written code never executes on the worker host, and Modal reclaims a sandbox at its own timeout whether or not the worker survives. Two contract details are worth knowing. read_file deliberately stays on the base class's shell implementation, because Modal's native read takes no length and so cannot honor a read budget: a read of /dev/zero never returns while stat reports it as zero-length. write_file and list_directory do use the native API. And a SandboxSpec naming allow_egress_to is refused unless egress_enforcement="sni", because Modal matches hostnames in the TLS handshake and leaves DNS open, which is weaker than the spec reads.
The provider verifier imports every submodule of a distribution directly, so it reached the backend's module-level `import modal` without the extra installed and failed with ModuleNotFoundError. The package's lazy `__getattr__` never runs on that path. Guard the import and raise AirflowOptionalProviderFeatureException, which the verifier ignores and which is what a Dag author reaching the module directly should see. The wrapper in the package `__getattr__` is dropped: the module now raises the right error itself, and catching ImportError there would also swallow a genuine one raised from inside the module. Also build the test fake's module object with setattr, since assigning attributes onto a ModuleType does not type-check, and drop a docs reference to an example Dag that does not exist in the provider.
kaxil
force-pushed
the
modal-sandbox-backend
branch
from
September 15, 2026 00:52
7e0219a to
4d783bc
Compare
The sandbox docs described the boundary but not the two things every author hits first: what to do about credentials, and how to get a result back. They also pointed at `KubernetesPodOperator` without saying why a pod is not the answer when an agent is choosing the commands. Add a worked example that drives a backend from a `@task`, which is the route that solves both: the credential is read from a connection at run time rather than fixed when the Dag file is parsed, and the artifact returns through the worker rather than through the model's context. Record two behaviours that were previously only visible in the code. On `sbx`, `SandboxSpec.env` is appended to `/etc/profile` inside the guest, so an injected value is readable by everything in that sandbox for its whole life, where the Modal backend passes the environment at creation instead. And the tools are text-only in both directions, so binary cannot round-trip through an agent even well under the byte caps.
The sandbox docs opened with the four tools and the boundary table, which answers "what does this do" for someone who already knows they want it, and leaves everyone else guessing. The question people actually arrive with is why this exists next to the narrower toolsets, since those run on the worker too. State the point directly: the toolset does not give an agent a new power, it moves a power the agent already had off the worker. Both `code_mode` and an agent skill's `run_skill_script` already execute model-written code in the worker process, and both say so. What changes here is where that code runs, not whether it can run at all. Follow it with the choice a reader is really making, including the two rows that send them somewhere else: name the operations with a narrow toolset when they can be named, and drive a backend from a task when the job is producing an artifact rather than deciding something.
…ming The new framing listed `code_mode` alongside an agent skill's `run_skill_script` as ways model-written code already runs on the worker, and said that in all of them the generated code sits beside the deployment's connections, filesystem and network position. That is true of a skill script and of a hand-rolled tool that shells out. It is not true of code mode: Monty confines generated code to the tools that were registered, which the boundary table a few paragraphs down already states. Describe code mode on its own terms instead. It is not a weaker sandbox, it is a different one, and it starts in well under a millisecond against roughly a second to provision a hosted sandbox. The reason to reach past it is capability rather than containment, because a restricted subset of Python with no shell, no package installation and no C extensions cannot do the work a model needs a real environment for.
Six facts decided whether an agent with a sandbox fits a Dag at all, and they were spread across four subsections and a parameter list. Three of them sat under the hosted backend's heading, so anyone evaluating with the local backend the docs tell you to develop against never read them. Two were not written down anywhere. Combining a sandbox with `durable=True` replays cached tool results describing a filesystem that no longer exists, and HITL regeneration starts a second agent run whose sandbox is empty while its history says otherwise. Neither raises, so both produce a confidently wrong answer, and the provider's front page sells durable step replay as the retry story without a word about it. State them together, at toolset level, including the two that send a reader elsewhere: an agent's sandbox cannot take a credential from a connection or a secrets backend, and an artifact leaves only through the model's context unless the sandboxed code writes it to storage itself. Also correct three smaller things a first reading tripped over: the whole-task row of the boundary table claimed it protects you from "everything" while the prose below explains it does not protect the task from its own code; the `image` parameter was documented as a registry tag when the code and the example both accept a prepared image; and the front page neither mentioned the toolset nor allowed for it when saying every tool runs in the worker. The artifact example returned the file's bytes, which pushes the artifact through XCom and stops working at the sizes the example exists to serve. It now writes to object storage and returns the location.
The code mode section explains that Monty cannot import third-party libraries and then lists the ways out, one of which was to reach for Docker or E2B through pydantic-ai-harness. That was the right advice before this provider had an answer of its own, and it is the wrong advice now: `SandboxToolset` solves exactly that problem, and the section never mentioned it. Every cross-link between the two features ran one direction. Recommend the toolset first, with the tradeoff stated in the terms that decide it: a real interpreter and installed packages against a sandbox per run and about a second to provision, where Monty starts in well under a millisecond. Also split the sentence describing how the two combine. It carried two claims at once and read as its own opposite, because the clause explaining why `run_command` stays a separate tool sat behind a "so" that appeared to describe the folded file tools. Say the two things separately, and answer the question it raised without meaning to: both call paths reach the same sandbox, because there is one per agent run however a call arrives.
… task Modal serves the native file operations from a helper binary it injects into the guest, and the SDK parses that helper's output. When the output is not what the SDK expects it raises out of `json`, not out of `modal.exception`, so `write_file` and `list_directory` let it past their handlers and the task died with a traceback the model never saw and could not act on. That breaks the contract the docs state, which is that only a terminal failure fails the task. Code running as root inside the sandbox can cause this deliberately by replacing the helper, and a helper that is broken or version-skewed causes it by accident. Translate it to a recoverable error naming `run_command`, which is a separate path that does not go through the helper, so the model has somewhere to go. Also refuse a single-label entry in `allow_egress_to`. A name such as `localhost` passed every check and then matched nothing, which is precisely the silent mismatch that validation exists to prevent, and the test suite asserted it was accepted while a neighbouring docstring explained why a name without dots allows nothing. Sharpen the egress documentation to match what the enforcement actually does. The destination address is not part of the decision at all: a connection opened to an unrelated address while presenting an allowed name is routed to the allowed host and answered by it, so the allowlist is a name-routed proxy rather than a bound on where packets may go. Name resolution is a two-way channel and not only an exfiltration route. Sharing a CDN is not by itself enough to defeat the allowlist, co-tenancy of the same TLS endpoint is. Record three guest-side facts a reader would otherwise assume wrongly: commands run as root, `workdir` is a starting directory rather than a jail, and `block_network=False` still leaves the metadata endpoint and private address ranges unreachable. Sandboxes also cannot reach each other.
A command shortened to fit the sandbox's remaining lifetime, which then uses all of it, has stood on the expiry. The backend asked a liveness probe what happened anyway, and at that exact instant the probe is not reliable: Modal reclaims a sandbox some seconds after its deadline, so the probe answers "alive" for a sandbox that dies moments later. Measured four runs out of four, the result came back with `timed_out` set and `sandbox_terminated` clear, so the model read only that its command was slow and the toolset kept a handle it should have dropped. The next tool call was then a coin flip, and three times out of three a call issued a few seconds later failed the task. Decide it from the clamp instead, which cannot race: if the deadline WAS the remaining lifetime and the command reached it, the lifetime is spent. The model is told the sandbox was replaced and the handle is dropped, which is the same handling a sandbox killed mid-command already got. An ordinary overrun still consults the sandbox, because there the command's budget and the sandbox's lifetime are different facts. Correct two claims that measurement contradicted. The `idle_timeout` note said a mid-run reclaim costs the agent its files; it costs the task, because the next call reaches a sandbox that is gone and that is terminal. And the egress section said non-TLS traffic to a listed host is refused after about thirty seconds. Nothing refuses it: the packets are dropped, one address stalls for around two minutes of TCP retries, and a client that tries every address a name resolves to multiplies that. Under the default sixty second command budget the model learns its command was slow rather than that the network stopped it. Document what follows from all of this for a Dag author: a run that outlives its sandbox fails the task, and the clock covers model generation and review pauses, so size `sandbox_timeout` against the whole run rather than against the commands in it.
The model reads tool descriptions and little else, and the sandbox's descriptions never mentioned egress, so the policy had to be discovered by failing. Under the default deny that costs a turn and returns a DNS error. Under an allowlist it is worse: nothing refuses plain HTTP, so a reach for `http://` burns the whole command budget and comes back as a timeout, which reads as "my command was slow" rather than "the network stopped me". The toolset already holds the spec, so it can say which of the three cases applies instead of hedging. A denied sandbox says so and names DNS, an allowlisted one lists the hosts and warns that plain HTTP fails anyway, and an open one says it has access. Also point at the working directory. Relative paths resolve against it, nothing named it, and under code mode there is no `pwd` among the folded callables, so a model had to fall back to the shell to find out where it was standing.
The tag query is documented as the way to find what a Dag left behind, which sends an operator looking at Modal for leaked sandboxes. Two of the surfaces they will reach for lie: the dashboard lists stopped sandboxes alongside live ones, and the `Tasks` count in `modal app list` lags by up to about a minute. Both showed sandboxes that had already been terminated, and in one case sent someone hunting an orphan that did not exist. Name the two that are reliable, `Sandbox.list` and `poll()`, and say what an integer from `poll()` means, since `None` for running and an exit status for stopped is the opposite way round from what people guess.
kaxil
force-pushed
the
modal-sandbox-backend
branch
from
September 15, 2026 02:50
08d98c4 to
5e76139
Compare
…d when a sandbox is billed
…eason not to use an agent
kaxil
marked this pull request as ready for review
September 20, 2026 22:19
Contributor
|
Quickest fix: git fetch upstream main && git rebase upstream/main
rm uv.lock && uv lock
git add uv.lock && git rebase --continue
git push --force-with-leaseAutomated nudge — ignore if you're not ready to rebase. This comment is updated in place on future |
vatsrahul1001
approved these changes
Sep 21, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
SandboxToolsetshipped in #68847 with one backend,sbx, which needs thesbxbinary on the worker host, a Docker login, a one-timesbx policy init, and KVM on Linux. An unprivileged container cannot provide the last of those, so every Kubernetes deployment has been unable to turn the toolset on at all, and its own docs say to run something else in production without offering one.ModalSandboxBackendis that something else. Each sandbox is a Modal sandbox provisioned over the API: nothing is installed on the worker, model-written code never executes on the worker host, and Modal reclaims a sandbox at its ownsandbox_timeoutwhether or not the worker survives.sbxhas no equivalent, so a dead worker leaves its microVM running.Credentials are ambient, as they are for Modal's own CLI:
modal token new, orMODAL_TOKEN_IDandMODAL_TOKEN_SECRETon the worker. Nothing is read until the first sandbox is created, so a Dag file that constructs the backend parses without them.What a sandbox is for
Adding this toolset gives the agent shell and file operations in a separate workspace. Every other tool keeps its existing permissions and runs where it ran before.
Whether that is a new capability depends on what the agent already had. An agent skill can ship a
run_skill_scripttool that runs a script on the worker, and the guidance there is already to switch it off when the worker holds anything sensitive. Teams without skills hand an agent a tool of their own that shells out. In both, the generated code runs beside the deployment's connections, the worker's filesystem, and its network position, and this is where that code should run instead. An agent that only had named database operations is being granted a general shell for the first time, in a workspace with less authority than the worker, and the author is choosing to grant it.code_modeis deliberately not in that list. Monty confines generated code to the tools you registered, so it is not arbitrary execution on the worker, and it starts in well under a millisecond against roughly a second to provision a hosted sandbox. You reach past it for a sandbox when the model needs something the interpreter does not have: a restricted subset of Python with no shell, no package installation and no C extensions cannot run the work a model needs a real environment for.So the question this answers is "where does the code the model writes run", and the answer covers only the four tools the toolset adds. Nothing else moves.
That also makes it the wrong tool whenever the work can be written down as a fixed set of operations. If you can name them, name them:
SQLToolsetandHookToolsetare narrower, bound their own results, and never expose a credential to the model, because the hook runs in the worker and only the result reaches the context.SQLToolset/HookToolsetcode_mode, which is faster and confines the glue to your registered toolsSandboxToolset@taskdriving a backend directly. That is a missing seam in the toolset rather than a recommendation, and it is tracked separatelyThe clearest case is the last of those third-row clauses, an agent that debugs its own code: it has to run something, read the real error and try again, and that loop cannot be enumerated in advance because each step depends on the previous step's output.
What it looks like
The four tools driven end to end in one sandbox, with a deliberate two second timeout in
the middle and the file from before it still readable afterwards, which is the behaviour
sbxcannot offer:And the boundary itself: the sandbox comes up with
block_network=True, carries none ofAirflow's environment, cannot resolve a hostname, and is destroyed when the task ends.
Design rationale
read_filedeliberately stays on the base class's shell implementation. Modal's native read takes no length argument, so honoring a read budget through it meansstatand then read: the TOCTOU windowbase.pydocuments avoiding, and no protection at all against anythingstatreports as zero-length. Measured againstmodal==1.5.5,read_bytes("/dev/zero")never returned in 195 seconds whilestatcalled that devicesize=0, and a 200 MB regular file came back whole, so there is no server-side ceiling either. The inherited implementation caps inside the guest withhead -c, so the size the caller trusts is the size that was actually transferred.write_fileandlist_directorydo use the native API, where it is a clear win.A
SandboxSpecnamingallow_egress_tois refused unless you passegress_enforcement="sni". Modal cannot combine an allowlist withblock_networkat all, and its hostname allowlist is matched against the name in the TLS handshake, which means TLS on 443 only, DNS left open for every hostname, and a host sharing a TLS endpoint with an allowed one still reachable. Measured: withpypi.orgas the only allowed host, a TLS session opened tofiles.pythonhosted.orgwhile presentingpypi.orgas the handshake name was allowed through and answered.block_network=Trueon its own maps exactly and takes DNS down with everything else, and it is the default. This mirrorssbx, which refuses an allowlist until the Deployment Manager declares the host policy.A command timeout does not cost you the sandbox. Modal stops the command server-side, so files written by earlier calls survive and the model can inspect them to work out what went wrong.
sbxhas to destroy the sandbox to be certain a command stopped.A timeout is a returncode rather than an exception, and not always the same one.
_ContainerProcess.polland.waitcatchExecTimeoutErrorinternally and set-1, carrying the SDK's own note that it should probably raise. The same command at the same budget has also been seen returning137, at the same elapsed time and from the same SDK version, so classification uses elapsed time rather than the status alone: a deadline lands at or past its budget while a guest killing itself lands nowhere near it.Two changes outside the backend
SandboxExecResultgains an optionalapplied_timeout, and the toolset reports a replaced sandbox after any failure rather than only after a timeout. Both exist because this backend produces statessbxnever did: it can shorten a command's deadline to fit what is left of a sandbox's life, and it can lose a sandbox under a command that looks like an ordinary non-zero exit. Without them the model is told a budget it did not have, or is handed a fresh empty sandbox with no explanation. Both fields have defaults, sosbxbehaviour does not change.Tradeoffs and limitations
Work does not survive the run. The sandbox is created on the model's first tool call and destroyed when the run ends, so an artifact the agent built leaves only through the model's context, which is bounded. If a task has to produce a file, drive a backend directly from a
@taskrather than handing it to an agent. This is a property of the toolset rather than of this backend, and the docs now say so plainly.A default sandbox cannot install packages, because the default spec denies all egress including DNS. Either bake what the agent needs into the image with a prepared
modal.Image, which needs no egress, or allowpypi.organdfiles.pythonhosted.organd accept the SNI caveats. Both are documented.One sandbox spans a whole agent run, so you pay for the model's thinking time between tool calls, not only for the seconds your commands run. Human output review is not part of that: it starts after
agent.run_synchas returned and the sandbox has been destroyed, so it costs no sandbox time and keeps no files. A default sandbox is about two cents an hour; a data-sized one is about two dollars.memoryis a request, not a ceiling. A sandbox created withmemory=512allocated 1.5 GiB without complaint, so it schedules the sandbox somewhere with room and does not bound what model-written code can take.Tags identify the Dag, not the run. They are fixed when the backend is constructed, and
createreceives no task context, so every sandbox in a mapped task carries the same ones.Known issues, not addressed here
durable=Trueand HITL regeneration both interact badly with any sandbox backend,sbxincluded: durable replays cached tool results describing a filesystem that no longer exists, and a regeneration hands the agent a fresh empty sandbox while its history insists the files are there. Neither raises, so the failure is a confidently wrong answer rather than an error.Both are pre-existing, and enforcing them is tracked separately rather than widened into this change. What this PR does add is the documentation, under a new toolset-level Limitations section, because the provider's own front page sells durable step replay as the retry story and nothing warned a reader away from the combination. That section also collects the other facts that decide whether a design works at all, which were previously spread across four subsections with three of them under a backend-specific heading where anyone evaluating
sbxwould never see them.