Skip to content

Fix XCom loss in deferrable KubernetesPodOperator when pod is already terminal - #72821

Draft
shepherd44 wants to merge 2 commits into
apache:mainfrom
shepherd44:fix/kpo-deferrable-xcom-shortcut
Draft

shepherd44 wants to merge 2 commits into
apache:mainfrom
shepherd44:fix/kpo-deferrable-xcom-shortcut

Conversation

@shepherd44

@shepherd44 shepherd44 commented Sep 9, 2026 •

Copy link
Copy Markdown

Why

When the pod reaches a terminal state before the operator defers, invoke_defer_method calls trigger_reentry inline instead of deferring:

if context and (
    pod_container_state == ContainerState.TERMINATED or pod_container_state == ContainerState.FAILED
):
    self.log.info("Skipping deferral as pod is already in a terminal state")
    self.trigger_reentry(...)          # return value dropped
else:
    self.defer(trigger=trigger, method_name="trigger_reentry", timeout=defer_timeout)

trigger_reentry ends with:

if self.do_xcom_push:
    return xcom_sidecar_output

That value was dropped at every level above it — invoke_defer_method, execute_async, and the deferrable branch of execute all called down without returning.

In the regular deferral path Airflow takes the return value of the resume method (trigger_reentry) and stores it as return_value, so the loss only happens on this shortcut. execute_sync already returns its result, which is why non-deferrable runs are unaffected.

The failure is silent. The task still succeeds, and pod_name / pod_namespace are pushed separately by execute_async, so the XCom entries look populated:

XCom keys
shortcut path pod_name, pod_namespace
regular deferral pod_name, pod_namespace, return_value

Downstream tasks pulling that XCom get None. In our deployment a Jinja template doing {{ ti.xcom_pull(task_ids="fetch")["data_key"] }} failed with 'None' has no attribute 'data_key', three retries deep, while the upstream task was reported as success. Retrying cannot help — the XCom is already gone.

Task logs make the two paths easy to tell apart:

# affected
Reusing existing pod '...' (phase=Running, reason=) since it is not terminated or evicted.
Skipping deferral as pod is already in a terminal state
Deleting pod: ...
                          <- no "Pushing xcom"

# working
Pausing task as DEFERRED.
Deleting pod: ...
Pushing xcom

Pods that finish within a second or two hit this often, which is why it showed up intermittently across several DAGs rather than as a hard failure.

What

Propagate the return value through the call sites, and widen the two return annotations (-> None → -> Any) accordingly.

SparkKubernetesOperator.execute dropped the deferrable result the same way, so it is fixed in the same commit — otherwise the fix would stop at KubernetesPodOperator and Spark tasks would still lose return_value.

New tests:

  • test_invoke_defer_method_returns_trigger_reentry_result_when_pod_already_terminal — the inline call's result is returned
  • test_execute_returns_deferrable_result — execute hands the deferrable result back, as the synchronous branch already does

Both fail on main with assert None == {'key': 'value'} and pass with the change.

One existing test was updated: test_execute_deferrable_does_not_call_super asserted result is None, which pinned the dropped-return behaviour rather than the "does not call super" contract the test is named for. It now asserts the result reaches the caller.

Scope notes

Testing

pytest providers/cncf/kubernetes/tests/unit/cncf/kubernetes/operators/test_pod.py \
       providers/cncf/kubernetes/tests/unit/cncf/kubernetes/operators/test_spark_kubernetes.py

300 passed. Two failures and ten errors (TestSuppress, write_logs) reproduce identically on unmodified main on my machine, so they are local environment issues rather than regressions from this change.

ruff check and ruff format --check pass on all touched files.

Gen-AI disclosure

This PR was prepared with the assistance of Gen-AI tools. The root cause was located by reading provider source and correlating it against Airflow API data and task logs from a real deployment where the bug occurred. I reviewed the diff and the tests, ran them locally, and I am able to explain and stand behind the change.

… terminal

When the pod reaches a terminal state before the operator defers,
`invoke_defer_method` calls `trigger_reentry` inline instead of deferring.
`trigger_reentry` returns the XCom sidecar output, but that value was
dropped: neither `invoke_defer_method`, nor `execute_async`, nor the
deferrable branch of `execute` returned it.

In the regular deferral path Airflow takes the return value of the resume
method (`trigger_reentry`) and stores it as `return_value`, so the loss only
happens on this shortcut. `execute_sync` already returns its result, which is
why non-deferrable runs are unaffected.

The task still succeeds and `pod_name` / `pod_namespace` are pushed by
`execute_async`, so the failure is silent: `do_xcom_push=True` produces no
`return_value` and downstream tasks pulling that XCom get `None`.

Pods that finish within a second or two hit this often.

Propagate the return value through the three call sites and widen the two
return annotations accordingly.
@boring-cyborg

boring-cyborg Bot commented Sep 9, 2026

Copy link
Copy Markdown

Congratulations on your first Pull Request and welcome to the Apache Airflow community! If you have any issues or are unsure about any anything please check our Contributors' Guide
Here are some useful points:

  • Pay attention to the quality of your code (ruff, mypy and type annotations). Our prek-hooks will help you with that.
  • In case of a new feature add useful documentation (in docstrings or in docs/ directory). Adding a new operator? Check this short guide Consider adding an example Dag that shows how users should use it.
  • Consider using Breeze environment for testing locally, it's a heavy docker but it ships with a working Airflow and a lot of integrations.
  • Be patient and persistent. It might take some time to get a review or get the final approval from Committers.
  • Please follow ASF Code of Conduct for all communication including (but not limited to) comments on Pull Requests, Mailing list and Slack.
  • Be sure to read the Airflow Coding style.
  • Always keep your Pull Requests rebased, otherwise your build might fail due to changes not related to your commits.
    Apache Airflow is a community-driven project and together we are making it better 🚀.
    In case of doubts contact the developers at:
    Mailing List: dev@airflow.apache.org
    Slack: https://s.apache.org/airflow-slack

@boring-cyborg boring-cyborg Bot added area:providers provider:cncf-kubernetes Kubernetes (k8s) provider related issues labels Sep 9, 2026
`SparkKubernetesOperator.execute` dropped the deferrable result the same way,
so the XCom fix would have stopped at `KubernetesPodOperator` and Spark tasks
would still lose `return_value`.

`test_execute_deferrable_does_not_call_super` asserted `result is None`, which
pinned the dropped-return behaviour rather than the "does not call super"
contract the test is named for. Updated to assert the result reaches the caller.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:providers provider:cncf-kubernetes Kubernetes (k8s) provider related issues

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant