Skip to content

Add a Modal backend for the sandbox toolset - #72910

Merged
kaxil merged 17 commits into
apache:mainfrom
astronomer:modal-sandbox-backend
Sep 21, 2026
Merged

kaxil merged 17 commits into
apache:mainfrom
astronomer:modal-sandbox-backend

Conversation

@kaxil

@kaxil kaxil commented Sep 10, 2026 •

Copy link
Copy Markdown
Member

SandboxToolset shipped in #68847 with one backend, sbx, which needs the sbx binary on the worker host, a Docker login, a one-time sbx policy init, and KVM on Linux. An unprivileged container cannot provide the last of those, so every Kubernetes deployment has been unable to turn the toolset on at all, and its own docs say to run something else in production without offering one.

ModalSandboxBackend is that something else. Each sandbox is a Modal sandbox provisioned over the API: nothing is installed on the worker, model-written code never executes on the worker host, and Modal reclaims a sandbox at its own sandbox_timeout whether or not the worker survives. sbx has no equivalent, so a dead worker leaves its microVM running.

from airflow.providers.common.ai.operators.agent import AgentOperator
from airflow.providers.common.ai.sandbox import ModalSandboxBackend
from airflow.providers.common.ai.toolsets import SandboxToolset

AgentOperator(
    task_id="sandboxed_analyst",
    prompt="Estimate pi with a Monte Carlo simulation of one million points.",
    llm_conn_id="pydanticai_default",
    toolsets=[SandboxToolset(ModalSandboxBackend())],
)

Credentials are ambient, as they are for Modal's own CLI: modal token new, or MODAL_TOKEN_ID and MODAL_TOKEN_SECRET on the worker. Nothing is read until the first sandbox is created, so a Dag file that constructs the backend parses without them.

What a sandbox is for

Adding this toolset gives the agent shell and file operations in a separate workspace. Every other tool keeps its existing permissions and runs where it ran before.

Whether that is a new capability depends on what the agent already had. An agent skill can ship a run_skill_script tool that runs a script on the worker, and the guidance there is already to switch it off when the worker holds anything sensitive. Teams without skills hand an agent a tool of their own that shells out. In both, the generated code runs beside the deployment's connections, the worker's filesystem, and its network position, and this is where that code should run instead. An agent that only had named database operations is being granted a general shell for the first time, in a workspace with less authority than the worker, and the author is choosing to grant it.

code_mode is deliberately not in that list. Monty confines generated code to the tools you registered, so it is not arbitrary execution on the worker, and it starts in well under a millisecond against roughly a second to provision a hosted sandbox. You reach past it for a sandbox when the model needs something the interpreter does not have: a restricted subset of Python with no shell, no package installation and no C extensions cannot run the work a model needs a real environment for.

So the question this answers is "where does the code the model writes run", and the answer covers only the four tools the toolset adds. Nothing else moves.

That also makes it the wrong tool whenever the work can be written down as a fixed set of operations. If you can name them, name them: SQLToolset and HookToolset are narrower, bound their own results, and never expose a credential to the model, because the hook runs in the worker and only the result reaches the context.

The agent needs to Reach for
Call operations you can name in advance SQLToolset / HookToolset
Chain those tools with glue logic code_mode, which is faster and confines the glue to your registered tools
Write and run open-ended code: reshape data with no known schema, install a package, fix its own failing script by reading the traceback SandboxToolset
Produce a large artifact for a downstream task Today, a @task driving a backend directly. That is a missing seam in the toolset rather than a recommendation, and it is tracked separately

The clearest case is the last of those third-row clauses, an agent that debugs its own code: it has to run something, read the real error and try again, and that loop cannot be enumerated in advance because each step depends on the previous step's output.

What it looks like

The four tools driven end to end in one sandbox, with a deliberate two second timeout in
the middle and the file from before it still readable afterwards, which is the behaviour
sbx cannot offer:

Task log showing six sandbox tool calls, a timeout, and recovery

And the boundary itself: the sandbox comes up with block_network=True, carries none of
Airflow's environment, cannot resolve a hostname, and is destroyed when the task ends.

Task log showing no Airflow environment inside the sandbox and DNS blocked

Design rationale

read_file deliberately stays on the base class's shell implementation. Modal's native read takes no length argument, so honoring a read budget through it means stat and then read: the TOCTOU window base.py documents avoiding, and no protection at all against anything stat reports as zero-length. Measured against modal==1.5.5, read_bytes("/dev/zero") never returned in 195 seconds while stat called that device size=0, and a 200 MB regular file came back whole, so there is no server-side ceiling either. The inherited implementation caps inside the guest with head -c, so the size the caller trusts is the size that was actually transferred. write_file and list_directory do use the native API, where it is a clear win.

A SandboxSpec naming allow_egress_to is refused unless you pass egress_enforcement="sni". Modal cannot combine an allowlist with block_network at all, and its hostname allowlist is matched against the name in the TLS handshake, which means TLS on 443 only, DNS left open for every hostname, and a host sharing a TLS endpoint with an allowed one still reachable. Measured: with pypi.org as the only allowed host, a TLS session opened to files.pythonhosted.org while presenting pypi.org as the handshake name was allowed through and answered. block_network=True on its own maps exactly and takes DNS down with everything else, and it is the default. This mirrors sbx, which refuses an allowlist until the Deployment Manager declares the host policy.

A command timeout does not cost you the sandbox. Modal stops the command server-side, so files written by earlier calls survive and the model can inspect them to work out what went wrong. sbx has to destroy the sandbox to be certain a command stopped.

A timeout is a returncode rather than an exception, and not always the same one. _ContainerProcess.poll and .wait catch ExecTimeoutError internally and set -1, carrying the SDK's own note that it should probably raise. The same command at the same budget has also been seen returning 137, at the same elapsed time and from the same SDK version, so classification uses elapsed time rather than the status alone: a deadline lands at or past its budget while a guest killing itself lands nowhere near it.

Two changes outside the backend

SandboxExecResult gains an optional applied_timeout, and the toolset reports a replaced sandbox after any failure rather than only after a timeout. Both exist because this backend produces states sbx never did: it can shorten a command's deadline to fit what is left of a sandbox's life, and it can lose a sandbox under a command that looks like an ordinary non-zero exit. Without them the model is told a budget it did not have, or is handed a fresh empty sandbox with no explanation. Both fields have defaults, so sbx behaviour does not change.

Tradeoffs and limitations

Work does not survive the run. The sandbox is created on the model's first tool call and destroyed when the run ends, so an artifact the agent built leaves only through the model's context, which is bounded. If a task has to produce a file, drive a backend directly from a @task rather than handing it to an agent. This is a property of the toolset rather than of this backend, and the docs now say so plainly.

A default sandbox cannot install packages, because the default spec denies all egress including DNS. Either bake what the agent needs into the image with a prepared modal.Image, which needs no egress, or allow pypi.org and files.pythonhosted.org and accept the SNI caveats. Both are documented.

One sandbox spans a whole agent run, so you pay for the model's thinking time between tool calls, not only for the seconds your commands run. Human output review is not part of that: it starts after agent.run_sync has returned and the sandbox has been destroyed, so it costs no sandbox time and keeps no files. A default sandbox is about two cents an hour; a data-sized one is about two dollars.

memory is a request, not a ceiling. A sandbox created with memory=512 allocated 1.5 GiB without complaint, so it schedules the sandbox somewhere with room and does not bound what model-written code can take.

Tags identify the Dag, not the run. They are fixed when the backend is constructed, and create receives no task context, so every sandbox in a mapped task carries the same ones.

Known issues, not addressed here

durable=True and HITL regeneration both interact badly with any sandbox backend, sbx included: durable replays cached tool results describing a filesystem that no longer exists, and a regeneration hands the agent a fresh empty sandbox while its history insists the files are there. Neither raises, so the failure is a confidently wrong answer rather than an error.

Both are pre-existing, and enforcing them is tracked separately rather than widened into this change. What this PR does add is the documentation, under a new toolset-level Limitations section, because the provider's own front page sells durable step replay as the retry story and nothing warned a reader away from the combination. That section also collects the other facts that decide whether a design works at all, which were previously spread across four subsections with three of them under a backend-specific heading where anyone evaluating sbx would never see them.

SandboxToolset has shipped with one backend, sbx, which needs the sbx binary on
the worker host and KVM on Linux. That rules out an unprivileged container, so
every Kubernetes deployment has been unable to use the toolset at all.

ModalSandboxBackend runs each sandbox in Modal instead. Nothing is installed on
the worker, model-written code never executes on the worker host, and Modal
reclaims a sandbox at its own timeout whether or not the worker survives.

Two contract details are worth knowing. read_file deliberately stays on the base
class's shell implementation, because Modal's native read takes no length and so
cannot honor a read budget: a read of /dev/zero never returns while stat reports
it as zero-length. write_file and list_directory do use the native API. And a
SandboxSpec naming allow_egress_to is refused unless egress_enforcement="sni",
because Modal matches hostnames in the TLS handshake and leaves DNS open, which
is weaker than the spec reads.
The provider verifier imports every submodule of a distribution directly, so it
reached the backend's module-level `import modal` without the extra installed and
failed with ModuleNotFoundError. The package's lazy `__getattr__` never runs on
that path.

Guard the import and raise AirflowOptionalProviderFeatureException, which the
verifier ignores and which is what a Dag author reaching the module directly
should see. The wrapper in the package `__getattr__` is dropped: the module now
raises the right error itself, and catching ImportError there would also swallow
a genuine one raised from inside the module.

Also build the test fake's module object with setattr, since assigning attributes
onto a ModuleType does not type-check, and drop a docs reference to an example Dag
that does not exist in the provider.
@kaxil
kaxil force-pushed the modal-sandbox-backend branch from 7e0219a to 4d783bc Compare September 15, 2026 00:52
The sandbox docs described the boundary but not the two things every author
hits first: what to do about credentials, and how to get a result back. They
also pointed at `KubernetesPodOperator` without saying why a pod is not the
answer when an agent is choosing the commands.

Add a worked example that drives a backend from a `@task`, which is the route
that solves both: the credential is read from a connection at run time rather
than fixed when the Dag file is parsed, and the artifact returns through the
worker rather than through the model's context.

Record two behaviours that were previously only visible in the code. On `sbx`,
`SandboxSpec.env` is appended to `/etc/profile` inside the guest, so an injected
value is readable by everything in that sandbox for its whole life, where the
Modal backend passes the environment at creation instead. And the tools are
text-only in both directions, so binary cannot round-trip through an agent even
well under the byte caps.
The sandbox docs opened with the four tools and the boundary table, which
answers "what does this do" for someone who already knows they want it, and
leaves everyone else guessing. The question people actually arrive with is why
this exists next to the narrower toolsets, since those run on the worker too.

State the point directly: the toolset does not give an agent a new power, it
moves a power the agent already had off the worker. Both `code_mode` and an
agent skill's `run_skill_script` already execute model-written code in the
worker process, and both say so. What changes here is where that code runs, not
whether it can run at all.

Follow it with the choice a reader is really making, including the two rows
that send them somewhere else: name the operations with a narrow toolset when
they can be named, and drive a backend from a task when the job is producing an
artifact rather than deciding something.
…ming

The new framing listed `code_mode` alongside an agent skill's
`run_skill_script` as ways model-written code already runs on the worker, and
said that in all of them the generated code sits beside the deployment's
connections, filesystem and network position. That is true of a skill script
and of a hand-rolled tool that shells out. It is not true of code mode: Monty
confines generated code to the tools that were registered, which the boundary
table a few paragraphs down already states.

Describe code mode on its own terms instead. It is not a weaker sandbox, it is
a different one, and it starts in well under a millisecond against roughly a
second to provision a hosted sandbox. The reason to reach past it is capability
rather than containment, because a restricted subset of Python with no shell,
no package installation and no C extensions cannot do the work a model needs a
real environment for.
Six facts decided whether an agent with a sandbox fits a Dag at all, and they
were spread across four subsections and a parameter list. Three of them sat
under the hosted backend's heading, so anyone evaluating with the local backend
the docs tell you to develop against never read them.

Two were not written down anywhere. Combining a sandbox with `durable=True`
replays cached tool results describing a filesystem that no longer exists, and
HITL regeneration starts a second agent run whose sandbox is empty while its
history says otherwise. Neither raises, so both produce a confidently wrong
answer, and the provider's front page sells durable step replay as the retry
story without a word about it.

State them together, at toolset level, including the two that send a reader
elsewhere: an agent's sandbox cannot take a credential from a connection or a
secrets backend, and an artifact leaves only through the model's context unless
the sandboxed code writes it to storage itself.

Also correct three smaller things a first reading tripped over: the whole-task
row of the boundary table claimed it protects you from "everything" while the
prose below explains it does not protect the task from its own code; the
`image` parameter was documented as a registry tag when the code and the
example both accept a prepared image; and the front page neither mentioned the
toolset nor allowed for it when saying every tool runs in the worker.

The artifact example returned the file's bytes, which pushes the artifact
through XCom and stops working at the sizes the example exists to serve. It now
writes to object storage and returns the location.
The code mode section explains that Monty cannot import third-party libraries
and then lists the ways out, one of which was to reach for Docker or E2B
through pydantic-ai-harness. That was the right advice before this provider had
an answer of its own, and it is the wrong advice now: `SandboxToolset` solves
exactly that problem, and the section never mentioned it. Every cross-link
between the two features ran one direction.

Recommend the toolset first, with the tradeoff stated in the terms that decide
it: a real interpreter and installed packages against a sandbox per run and
about a second to provision, where Monty starts in well under a millisecond.

Also split the sentence describing how the two combine. It carried two claims
at once and read as its own opposite, because the clause explaining why
`run_command` stays a separate tool sat behind a "so" that appeared to describe
the folded file tools. Say the two things separately, and answer the question it
raised without meaning to: both call paths reach the same sandbox, because
there is one per agent run however a call arrives.
… task

Modal serves the native file operations from a helper binary it injects into
the guest, and the SDK parses that helper's output. When the output is not what
the SDK expects it raises out of `json`, not out of `modal.exception`, so
`write_file` and `list_directory` let it past their handlers and the task died
with a traceback the model never saw and could not act on. That breaks the
contract the docs state, which is that only a terminal failure fails the task.

Code running as root inside the sandbox can cause this deliberately by
replacing the helper, and a helper that is broken or version-skewed causes it
by accident. Translate it to a recoverable error naming `run_command`, which is
a separate path that does not go through the helper, so the model has somewhere
to go.

Also refuse a single-label entry in `allow_egress_to`. A name such as
`localhost` passed every check and then matched nothing, which is precisely the
silent mismatch that validation exists to prevent, and the test suite asserted
it was accepted while a neighbouring docstring explained why a name without
dots allows nothing.

Sharpen the egress documentation to match what the enforcement actually does.
The destination address is not part of the decision at all: a connection opened
to an unrelated address while presenting an allowed name is routed to the
allowed host and answered by it, so the allowlist is a name-routed proxy rather
than a bound on where packets may go. Name resolution is a two-way channel and
not only an exfiltration route. Sharing a CDN is not by itself enough to defeat
the allowlist, co-tenancy of the same TLS endpoint is.

Record three guest-side facts a reader would otherwise assume wrongly: commands
run as root, `workdir` is a starting directory rather than a jail, and
`block_network=False` still leaves the metadata endpoint and private address
ranges unreachable. Sandboxes also cannot reach each other.
A command shortened to fit the sandbox's remaining lifetime, which then uses all
of it, has stood on the expiry. The backend asked a liveness probe what happened
anyway, and at that exact instant the probe is not reliable: Modal reclaims a
sandbox some seconds after its deadline, so the probe answers "alive" for a
sandbox that dies moments later. Measured four runs out of four, the result came
back with `timed_out` set and `sandbox_terminated` clear, so the model read only
that its command was slow and the toolset kept a handle it should have dropped.
The next tool call was then a coin flip, and three times out of three a call
issued a few seconds later failed the task.

Decide it from the clamp instead, which cannot race: if the deadline WAS the
remaining lifetime and the command reached it, the lifetime is spent. The model
is told the sandbox was replaced and the handle is dropped, which is the same
handling a sandbox killed mid-command already got. An ordinary overrun still
consults the sandbox, because there the command's budget and the sandbox's
lifetime are different facts.

Correct two claims that measurement contradicted. The `idle_timeout` note said a
mid-run reclaim costs the agent its files; it costs the task, because the next
call reaches a sandbox that is gone and that is terminal. And the egress section
said non-TLS traffic to a listed host is refused after about thirty seconds.
Nothing refuses it: the packets are dropped, one address stalls for around two
minutes of TCP retries, and a client that tries every address a name resolves to
multiplies that. Under the default sixty second command budget the model learns
its command was slow rather than that the network stopped it.

Document what follows from all of this for a Dag author: a run that outlives its
sandbox fails the task, and the clock covers model generation and review pauses,
so size `sandbox_timeout` against the whole run rather than against the commands
in it.
The model reads tool descriptions and little else, and the sandbox's
descriptions never mentioned egress, so the policy had to be discovered by
failing. Under the default deny that costs a turn and returns a DNS error.
Under an allowlist it is worse: nothing refuses plain HTTP, so a reach for
`http://` burns the whole command budget and comes back as a timeout, which
reads as "my command was slow" rather than "the network stopped me".

The toolset already holds the spec, so it can say which of the three cases
applies instead of hedging. A denied sandbox says so and names DNS, an
allowlisted one lists the hosts and warns that plain HTTP fails anyway, and an
open one says it has access.

Also point at the working directory. Relative paths resolve against it, nothing
named it, and under code mode there is no `pwd` among the folded callables, so a
model had to fall back to the shell to find out where it was standing.
The tag query is documented as the way to find what a Dag left behind, which
sends an operator looking at Modal for leaked sandboxes. Two of the surfaces
they will reach for lie: the dashboard lists stopped sandboxes alongside live
ones, and the `Tasks` count in `modal app list` lags by up to about a minute.
Both showed sandboxes that had already been terminated, and in one case sent
someone hunting an orphan that did not exist.

Name the two that are reliable, `Sandbox.list` and `poll()`, and say what an
integer from `poll()` means, since `None` for running and an exit status for
stopped is the opposite way round from what people guess.
@kaxil
kaxil force-pushed the modal-sandbox-backend branch from 08d98c4 to 5e76139 Compare September 15, 2026 02:50
@kaxil
kaxil marked this pull request as ready for review September 20, 2026 22:19
@github-actions

Copy link
Copy Markdown
Contributor

uv.lock on main just moved via #73363 ("Add support for TypeSafe Jev classifier models"), commit 134b733 and this PR currently conflicts.

Quickest fix:

git fetch upstream main && git rebase upstream/main
rm uv.lock && uv lock
git add uv.lock && git rebase --continue
git push --force-with-lease

Automated nudge — ignore if you're not ready to rebase. This comment is updated in place on future uv.lock bumps.

@vatsrahul1001 vatsrahul1001 left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

One nit, non-blocking: in _still_alive the bare except Exception returns True, so a non-Modal error there gets read as "alive" and the handle is reused. Intent seems right, just no test for that branch — might be worth one.

@kaxil
kaxil merged commit 51876b0 into apache:main Sep 21, 2026
97 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants