Repository navigation
Cache repeated agent prompts by default in AgentOperator - #73994
Merged
Merged
Conversation
vatsrahul1001
approved these changes
Oct 1, 2026
vatsrahul1001
left a comment
Contributor
There was a problem hiding this comment.
Looks good, Conflict needs to be resolved
AgentOperator(cache_prompt=True) turns on each provider's prompt caching: Anthropic, Bedrock Converse and OpenRouter get breakpoints on the tool definitions, system prompt and latest message; OpenAI and Gemini already cache on their own. A provider's own cache settings from the caller, the connection's model or a spec file take that provider over entirely.
…s more A CachePoint in the prompt or history now turns PromptCaching off, since Anthropic needs a longer-lived entry before a shorter one. The docs cover the final uncached tail write and map indexes that start together. The failure-path log gets the cache line too.
kaxil
force-pushed
the
commonai-cache-prompt
branch
from
October 1, 2026 10:49
06d8482 to
788f12a
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
An agent re-sends its tool definitions, system prompt and conversation on every request. A tool-calling agent makes several requests per task, and a mapped
@task.agentrepeats all of them per map index with the same system prompt. OpenAI and Gemini cache that repeated prefix on their own; Anthropic, Bedrock and OpenRouter cache only when each request is marked, and pydantic-ai spells the marks differently per provider (anthropic_cache_*,bedrock_cache_*,openrouter_cache_*inmodel_settings). In practice nobody sets them, so these agents pay full input price for the same prefix over and over.This adds
cache_prompt: bool = TruetoAgentOperator(and so@task.agent). It turns on each provider's own prompt caching for the tool definitions, the system prompt and the latest message, and is a no-op for providers that cache automatically. Setting itFalserestores the old requests exactly.Evidence
One run of a test Dag against
claude-opus-4-8, the first time that Dag ran, so nothing was cached beforehand. The system prompt is 60 lines of triage policy. Every task makes two requests (one tool call), each sending about 6,000 input tokens. The tasks ran in the order shown. Cache read and write counts are the usage Anthropic returned and are included in input tokens; costs are what pydantic-ai priced and the task logged.cache_prompttriage_firstTruetriage_mapped, map index 2Truetriage_mapped, map index 0Truetriage_mapped, map index 1Truetriage_uncachedFalsetriage_uncachedpays full price for all 12,160 input tokens.triage_firstwrites 6,136 tokens to the cache and reads 6,027 back on its second request. It costs 28% less thantriage_uncached; the saving is smaller than on the mapped tasks because writing to the cache costs 1.25x the input price.triage_uncached. Both of its requests read from the cache, since one request can read back at most the ~6,000 tokens of the request before it. The first request reads the tool definitions and system prompt that an earlier task wrote.Every cost matches what
genai-prices0.1.8, the pricing library pydantic-ai uses, gives from the logged token counts. Itsclaude-opus-4-8rates are $5/M input, $25/M output, $0.50/M cache read and $6.25/M cache write. For map index 0: 4 × $5/M + 197 × $6.25/M + 11,942 × $0.50/M + 105 × $25/M = $0.00984725, the logged figure.Two more tasks in the same run check behaviour rather than savings:
caller_automaticsetsanthropic_cache=Trueitself and succeeds. pydantic-ai refuses a request that combines that setting withanthropic_cache_messages, so the flag did not add its own.triage_openairunsgpt-5.4-miniwith the flag on and succeeds. OpenAI reported 3,584 cached input tokens from its own automatic caching; the flag sets nothing for OpenAI.Task log lines from the run
The same lines in the task log view, from a second run of the Dag a few minutes later, while the first run's cache was still live:
Design rationale
A pydantic-ai capability, not a settings merge in the operator.
PromptCaching.get_model_settings()returns a callable, which pydantic-ai calls with the settings merged so far: the model's own, the agent'smodel_settings(static or callable), and a spec file's. It fills in only the providers none of those configure. Merging a dict intoagent_params["model_settings"]could not see a spec file or a callable, and would have had to guess who wins.Anything the caller set for a provider takes that provider over entirely. Setting any
anthropic_cache*key means nothing is added for Anthropic, rather than filling in the keys that are missing. pydantic-ai refusesanthropic_cachetogether withanthropic_cache_messages, so filling keys one by one would turn a valid caller setting into an error. ACachePointin the prompt or the message history turns all of it off, because Anthropic requires a longer-lived cache entry to come before a shorter one, and our 5-minute marks on the tools and system prompt would land ahead of a caller's 1-hour mark.anthropic_cache_messagesrather than the top-levelanthropic_cache. Both move the cache mark to the latest message on each request. pydantic-ai documents the per-block form as the one that works with Anthropic-compatible gateways that lack the top-level parameter, and on Bedrock and Vertexanthropic_cachefalls back to it anyway. A caller who prefersanthropic_cachesets it and gets it, as the table shows.No provider detection. Each pydantic-ai model reads only its own provider's settings and ignores the rest, and Bedrock and OpenRouter add the marks only for models whose profile supports caching. One set of settings therefore covers a
FallbackModelchain that spans providers, with nothing resolved in the hook.Why on by default. The people who save the most (long system prompts, tool loops, fan-out over rows) are the ones least likely to find three provider-specific flags. A prompt shorter than Anthropic's minimum length for caching (512 to 4,096 tokens depending on the model) is not cached and costs nothing extra. The docs name the cases where caching costs more than it saves: a single long request never repeated within five minutes, a final tool result much larger than the prompt, and map indexes that all start at the same moment.
Cache settings are excluded from the durable-execution request fingerprint, so switching
cache_promptbetween attempts does not re-run steps an earlier attempt completed.Gotchas
input_tokensin the log and theusageXCom includes the cached tokens; the newLLM prompt cache:line is where the split shows. The XCom does not carry the split.AgentOperatorand@task.agentget the flag.LLMOperatorand the otherLLM*operators build agents the same way and could take it in a follow-up.{pr_number}.significant.rst, in airflow-core/newsfragments. You can add this file in a follow-up commit after the PR is created so you know the PR number.