Skip to content

Cache repeated agent prompts by default in AgentOperator - #73994

Merged
kaxil merged 4 commits into
apache:mainfrom
astronomer:commonai-cache-prompt
Oct 1, 2026
Merged

kaxil merged 4 commits into
apache:mainfrom
astronomer:commonai-cache-prompt

Conversation

@kaxil

@kaxil kaxil commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

An agent re-sends its tool definitions, system prompt and conversation on every request. A tool-calling agent makes several requests per task, and a mapped @task.agent repeats all of them per map index with the same system prompt. OpenAI and Gemini cache that repeated prefix on their own; Anthropic, Bedrock and OpenRouter cache only when each request is marked, and pydantic-ai spells the marks differently per provider (anthropic_cache_*, bedrock_cache_*, openrouter_cache_* in model_settings). In practice nobody sets them, so these agents pay full input price for the same prefix over and over.

This adds cache_prompt: bool = True to AgentOperator (and so @task.agent). It turns on each provider's own prompt caching for the tool definitions, the system prompt and the latest message, and is a no-op for providers that cache automatically. Setting it False restores the old requests exactly.

Evidence

One run of a test Dag against claude-opus-4-8, the first time that Dag ran, so nothing was cached beforehand. The system prompt is 60 lines of triage policy. Every task makes two requests (one tool call), each sending about 6,000 input tokens. The tasks ran in the order shown. Cache read and write counts are the usage Anthropic returned and are included in input tokens; costs are what pydantic-ai priced and the task logged.

Task cache_prompt Input tokens Cache read Cache write Output tokens Cost
triage_first True 12,167 6,027 6,136 205 $0.0465
triage_mapped, map index 2 True 12,141 11,941 196 106 $0.0099
triage_mapped, map index 0 True 12,143 11,942 197 105 $0.0098
triage_mapped, map index 1 True 12,156 11,947 205 169 $0.0115
triage_uncached False 12,160 0 0 150 $0.0646
  • Caching off: triage_uncached pays full price for all 12,160 input tokens.
  • First task: triage_first writes 6,136 tokens to the cache and reads 6,027 back on its second request. It costs 28% less than triage_uncached; the saving is smaller than on the mapped tasks because writing to the cache costs 1.25x the input price.
  • Mapped tasks: each map index reads about 11,940 of its roughly 12,150 input tokens from the cache and costs 5.6x to 6.6x less than triage_uncached. Both of its requests read from the cache, since one request can read back at most the ~6,000 tokens of the request before it. The first request reads the tool definitions and system prompt that an earlier task wrote.

Every cost matches what genai-prices 0.1.8, the pricing library pydantic-ai uses, gives from the logged token counts. Its claude-opus-4-8 rates are $5/M input, $25/M output, $0.50/M cache read and $6.25/M cache write. For map index 0: 4 × $5/M + 197 × $6.25/M + 11,942 × $0.50/M + 105 × $25/M = $0.00984725, the logged figure.

Two more tasks in the same run check behaviour rather than savings:

  • caller_automatic sets anthropic_cache=True itself and succeeds. pydantic-ai refuses a request that combines that setting with anthropic_cache_messages, so the flag did not add its own.
  • triage_openai runs gpt-5.4-mini with the flag on and succeeds. OpenAI reported 3,584 cached input tokens from its own automatic caching; the flag sets nothing for OpenAI.
Task log lines from the run
-- triage_first
LLM run complete: model=claude-opus-4-8, requests=2, tool_calls=1, input_tokens=12167, output_tokens=205, total_tokens=12372
LLM prompt cache: cache_read_tokens=6027, cache_write_tokens=6136
LLM run cost: $0.04650850 (USD, best-effort)
-- triage_mapped (map_index=2)
LLM run complete: model=claude-opus-4-8, requests=2, tool_calls=1, input_tokens=12141, output_tokens=106, total_tokens=12247
LLM prompt cache: cache_read_tokens=11941, cache_write_tokens=196
LLM run cost: $0.00986550 (USD, best-effort)
-- triage_mapped (map_index=0)
LLM run complete: model=claude-opus-4-8, requests=2, tool_calls=1, input_tokens=12143, output_tokens=105, total_tokens=12248
LLM prompt cache: cache_read_tokens=11942, cache_write_tokens=197
LLM run cost: $0.00984725 (USD, best-effort)
-- triage_mapped (map_index=1)
LLM run complete: model=claude-opus-4-8, requests=2, tool_calls=1, input_tokens=12156, output_tokens=169, total_tokens=12325
LLM prompt cache: cache_read_tokens=11947, cache_write_tokens=205
LLM run cost: $0.01149975 (USD, best-effort)
-- triage_uncached
LLM run complete: model=claude-opus-4-8, requests=2, tool_calls=1, input_tokens=12160, output_tokens=150, total_tokens=12310
LLM run cost: $0.064550 (USD, best-effort)
-- caller_automatic
LLM run complete: model=claude-opus-4-8, requests=2, tool_calls=1, input_tokens=12143, output_tokens=104, total_tokens=12247
LLM prompt cache: cache_read_tokens=12139, cache_write_tokens=0
LLM run cost: $0.0086895 (USD, best-effort)
-- triage_openai
LLM run complete: model=gpt-5.4-mini, requests=2, tool_calls=1, input_tokens=7890, output_tokens=36, total_tokens=7926
LLM prompt cache: cache_read_tokens=3584, cache_write_tokens=0
LLM run cost: $0.0036603 (USD, best-effort)

The same lines in the task log view, from a second run of the Dag a few minutes later, while the first run's cache was still live:

cache_prompt=False
cache_prompt=True
cache_prompt=True, map index 1

Design rationale

A pydantic-ai capability, not a settings merge in the operator. PromptCaching.get_model_settings() returns a callable, which pydantic-ai calls with the settings merged so far: the model's own, the agent's model_settings (static or callable), and a spec file's. It fills in only the providers none of those configure. Merging a dict into agent_params["model_settings"] could not see a spec file or a callable, and would have had to guess who wins.

Anything the caller set for a provider takes that provider over entirely. Setting any anthropic_cache* key means nothing is added for Anthropic, rather than filling in the keys that are missing. pydantic-ai refuses anthropic_cache together with anthropic_cache_messages, so filling keys one by one would turn a valid caller setting into an error. A CachePoint in the prompt or the message history turns all of it off, because Anthropic requires a longer-lived cache entry to come before a shorter one, and our 5-minute marks on the tools and system prompt would land ahead of a caller's 1-hour mark.

anthropic_cache_messages rather than the top-level anthropic_cache. Both move the cache mark to the latest message on each request. pydantic-ai documents the per-block form as the one that works with Anthropic-compatible gateways that lack the top-level parameter, and on Bedrock and Vertex anthropic_cache falls back to it anyway. A caller who prefers anthropic_cache sets it and gets it, as the table shows.

No provider detection. Each pydantic-ai model reads only its own provider's settings and ignores the rest, and Bedrock and OpenRouter add the marks only for models whose profile supports caching. One set of settings therefore covers a FallbackModel chain that spans providers, with nothing resolved in the hook.

Why on by default. The people who save the most (long system prompts, tool loops, fan-out over rows) are the ones least likely to find three provider-specific flags. A prompt shorter than Anthropic's minimum length for caching (512 to 4,096 tokens depending on the model) is not cached and costs nothing extra. The docs name the cases where caching costs more than it saves: a single long request never repeated within five minutes, a final tool result much larger than the prompt, and map indexes that all start at the same moment.

Cache settings are excluded from the durable-execution request fingerprint, so switching cache_prompt between attempts does not re-run steps an earlier attempt completed.

Gotchas

  • input_tokens in the log and the usage XCom includes the cached tokens; the new LLM prompt cache: line is where the split shows. The XCom does not carry the split.
  • Only AgentOperator and @task.agent get the flag. LLMOperator and the other LLM* operators build agents the same way and could take it in a follow-up.

  • Read the Pull Request Guidelines for more information. Note: commit author/co-author name and email in commits become permanently public when merged.
  • For fundamental code changes, an Airflow Improvement Proposal (AIP) is needed.
  • When adding dependency, check compliance with the ASF 3rd Party License Policy.
  • For significant user-facing changes create newsfragment: {pr_number}.significant.rst, in airflow-core/newsfragments. You can add this file in a follow-up commit after the PR is created so you know the PR number.

@vatsrahul1001 vatsrahul1001 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good, Conflict needs to be resolved

kaxil added 4 commits October 1, 2026 11:42
AgentOperator(cache_prompt=True) turns on each provider's prompt caching:
Anthropic, Bedrock Converse and OpenRouter get breakpoints on the tool
definitions, system prompt and latest message; OpenAI and Gemini already
cache on their own. A provider's own cache settings from the caller, the
connection's model or a spec file take that provider over entirely.
…s more

A CachePoint in the prompt or history now turns PromptCaching off, since
Anthropic needs a longer-lived entry before a shorter one. The docs cover
the final uncached tail write and map indexes that start together. The
failure-path log gets the cache line too.
@kaxil
kaxil force-pushed the commonai-cache-prompt branch from 06d8482 to 788f12a Compare October 1, 2026 10:49
@kaxil
kaxil merged commit 27cd88c into apache:main Oct 1, 2026
85 checks passed
@kaxil
kaxil deleted the commonai-cache-prompt branch October 1, 2026 11:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants