Skip to content

KubernetesExecutor: multi-scheduler completed-pod thrash (10.15.0+) #66396

Description

@potiuk

Apache Airflow Provider versions

apache-airflow-providers-cncf-kubernetes >= 10.15.0
(verified on main; introduced by #61839, released in 10.15.0)

Apache Airflow version

3.x with KubernetesExecutor and 2 or more schedulers.

What happened

PR #61839 added a periodic call to _adopt_completed_pods from inside
KubernetesExecutor.sync(), gated by
[scheduler] orphaned_tasks_check_interval (default 300 s):

https://github.com/apache/airflow/blob/main/providers/cncf/kubernetes/src/airflow/providers/cncf/kubernetes/executors/kubernetes_executor.py#L271-L275

_adopt_completed_pods selects every Succeeded pod whose airflow-worker
label is not the current scheduler's label and PATCHes it with the
current scheduler's label so its KubernetesJobWatcher will see the change
and DELETE the pod:

https://github.com/apache/airflow/blob/main/providers/cncf/kubernetes/src/airflow/providers/cncf/kubernetes/executors/kubernetes_executor.py#L684-L714

query_kwargs = {
    "field_selector": "status.phase=Succeeded",
    "label_selector": (
        "kubernetes_executor=True,"
        f"airflow-worker!={new_worker_id_label},{POD_EXECUTOR_DONE_KEY}!=True"
    ),
}

With N schedulers (N ≥ 2) running concurrently, every
orphaned_tasks_check_interval each scheduler iterates over every Succeeded
pod that doesn't carry its own label and PATCHes it. Schedulers fight
each other:

  • Scheduler A relabels every Succeeded pod owned by B and C → A's watcher
    receives the update event and starts DELETEing
  • Scheduler B does the same a few seconds later → relabels A's freshly
    patched pods to B → B's watcher takes over
  • Scheduler C the same

At steady state with high pod churn (we have a user reporting 1000+ pods
created in a couple of minutes — see context below), this manifests as:

  • Heavy PATCH /api/v1/namespaces/.../pods/... traffic against the kube
    API server, multiplied by N
  • Heavy _list_pods traffic — see Kubernetes Executor List Pods Performance Improvement #35599 for the known per-call cost at
    scale (15–30 s with 500 pods, returning the full V1PodList)
  • Tasks stalling in scheduled / queued because every scheduler loop
    is burning seconds inside _list_pods + patch_namespaced_pod and not
    doing useful scheduling
  • delete_worker_pods=False does NOT help — the periodic adoption code
    path doesn't gate on delete_worker_pods, it goes through the watcher's
    delete

What you think should happen instead

The periodic _adopt_completed_pods should only adopt pods owned by
schedulers that are no longer alive — i.e., scope the
airflow-worker!=<my_label> query to also exclude airflow-worker IN (<set of currently-active scheduler job IDs>). The set of active
scheduler job IDs is already known to the scheduler via the
SchedulerJob rows. With that scoping:

  • Single-scheduler deployment: no behavior change.
  • Multi-scheduler deployment: each scheduler only adopts orphaned pods
    whose owning scheduler is gone — no thrash, no cross-scheduler
    relabeling.

The original goal of #61839 (closing #57553 — completed pods leaking
after scheduler restart) is still achieved.

How to reproduce

  1. Deploy Airflow 3.x with KubernetesExecutor and at least 2
    schedulers using cncf.kubernetes >= 10.15.0.
  2. Run a workload that produces enough Succeeded pods to keep them
    visible across the 5-minute orphaned_tasks_check_interval window
    (anything sustained for > 5 min works).
  3. Observe PATCH /api/v1/namespaces/.../pods/... traffic against the
    kube API server — each Succeeded pod gets PATCHed by every scheduler
    that doesn't currently own its label, on every interval tick.
  4. The reproducer in the original PR (k8s executor - ensure pods cleaned up #61839, Completed Kubernetes Pods not cleared up #57553) still works for
    the single-scheduler "completed pods leak after restart" case, so
    any fix needs to keep that scenario working.

Operating system

N/A (the bug is in the Python code path, not OS-specific).

Versions of Apache Airflow Providers

  • apache-airflow-providers-cncf-kubernetes 10.15.0 — present on main.

Deployment

Official Apache Airflow Helm Chart and any other Kubernetes deployment
running multiple schedulers + KubernetesExecutor.

Anything else

  • Caught while triaging a user report on the Airflow user mailing list:
    3 schedulers, Airflow 3.2.1, "thrashing between the schedulers trying
    to PATCH/DELETE pods after task completion" with delete_worker_pods=False
    not helping. Their workaround was to move workloads to Celery, which
    bypasses this code path entirely.
  • Workaround for affected users until a fix lands: set
    [scheduler] orphaned_tasks_check_interval very high (e.g., 86400)
    and rely on a separate delete_worker_pods cron for cleanup.
  • Original closing-issue / introducing-PR / related-perf-issue:

Are you willing to submit PR?

  • Yes I am willing to submit a PR! (per the discussion in the user
    thread, I offered to draft a fix that scopes the periodic adoption
    to dead schedulers only)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:schedulerkind:bugThis is a clearly a bugpriority:highHigh priority bug that should be patched quickly but does not require immediate new releaseprovider:cncf-kubernetesKubernetes (k8s) provider related issues

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions