Skip to content

Stop failing scheduled tasks when a serialized Dag is briefly missing - #72814

Closed
rjgoyln wants to merge 1 commit into
apache:mainfrom
rjgoyln:repro/62050-serialized-dag-bulk-fail
Closed

rjgoyln wants to merge 1 commit into
apache:mainfrom
rjgoyln:repro/62050-serialized-dag-bulk-fail

Conversation

@rjgoyln

@rjgoyln rjgoyln commented Sep 9, 2026 •

Copy link
Copy Markdown
Contributor

Summary

When a run's serialized Dag cannot be resolved during the task-concurrency check, the scheduler currently fails every SCHEDULED task instance of that Dag in one unbounded UPDATE, across all runs.

This is problematic for several reasons:

  • The update bypasses handle_failure, so the affected task instances lose their normal retries, callbacks, and failure logs.
  • It can affect task instances belonging to other runs of the same Dag.
  • The DagRun itself remains RUNNING, leaving manual clearing as the only recovery.
  • The missing serialized Dag is often a transient inconsistency between the run's pinned dag_version and the serialized Dag row it resolves to.

The UPDATE also has a load-bearing side effect: moving the task instances to FAILED keeps them out of the next critical-section query. Simply removing the update would allow the same task instances to refill every batch and potentially starve other Dags.

Change

Instead of failing the task instances:

  • Keep them SCHEDULED and skip them for the current scheduler round.
  • Track the affected Dag runs as starved so subsequent queries can reach other Dags.
  • Log the resolution failure once per Dag per round rather than once per task instance.

The starvation key is (dag_id, run_id) rather than just dag_id because serialized-Dag resolution is performed per run. One run may fail to resolve while another run of the same Dag still resolves successfully.

Behavior change

A run whose serialized Dag remains unresolvable now leaves its task instances SCHEDULED instead of marking them FAILED.

This preserves the possibility of recovery if the serialized Dag becomes available again. The DagRun cannot progress while its serialized Dag remains unavailable, so the run was already unable to make progress before this change.

Trade-off

If a run never becomes resolvable, its task instances remain SCHEDULED and are re-evaluated in subsequent scheduler rounds rather than being removed from consideration by changing their state to FAILED.

The additional scheduling cost is limited to cases where these task instances are ahead of work that could otherwise be scheduled. When they fill a batch, the scheduler skips them and continues to the work behind them. They occupy no pool slots and do not count toward task-concurrency limits while they remain SCHEDULED.

There is currently no mechanism that terminates such a run: dagrun_timeout is evaluated only after the serialized Dag has been resolved, while task_queued_timeout only considers QUEUED task instances.

Handling the terminal lifecycle of a DagRun whose serialized Dag is permanently unavailable is outside the scope of this PR.

Tests

  • Four runs with missing serialized Dags and max_tis=2 still allow a healthy Dag to queue. The serialized Dag rows are deleted rather than mocked.
  • A missing serialized Dag for one run does not prevent another run of the same Dag from being scheduled.
  • A task instance left SCHEDULED after a transient miss is queued successfully on the next scheduler round once the serialized Dag becomes available again.

Relation to #72652

#72652 addresses the same starvation issue using a per-Dag starvation key.

This change instead tracks starvation per (dag_id, run_id), deduplicates the resolution error once per Dag per round, and adds regression coverage for the starvation and per-run behavior.

closes: #62050

Was generative AI tooling used to co-author this PR?
  • Yes — Claude Code (Opus 5)

Generated-by: Claude Code (Opus 5) following the guidelines

@rjgoyln
rjgoyln force-pushed the repro/62050-serialized-dag-bulk-fail branch from a94b5a9 to 495f5b3 Compare September 9, 2026 15:10
A serialized Dag that cannot be read is almost always transient -- the Dag processor
is mid-write, or the row was briefly unreadable -- but the scheduler treated it as
final and failed every SCHEDULED task instance of the Dag, across all of its Dag
runs, bypassing retries and failure callbacks. The Dag run was not finished either,
so those tasks could only be recovered by clearing them by hand.

The bulk update also doubled as a starvation filter: flipping the rows to FAILED is
what kept them out of the next iteration of the critical-section query. Starving the
run preserves that effect, so an unresolvable Dag cannot hold up the Dags queued
behind it. The error is reported once per Dag rather than once per task instance, so
a Dag with many stuck runs cannot flood the scheduler log now that its task
instances survive the round.

closes: apache#62050
@rjgoyln
rjgoyln force-pushed the repro/62050-serialized-dag-bulk-fail branch 2 times, most recently from 495f5b3 to dd58070 Compare September 11, 2026 14:06
@rjgoyln
rjgoyln marked this pull request as ready for review September 11, 2026 15:48
@potiuk potiuk added the closed because of open PR limit Closed as a one-time step of introducing the open pull request limit label Sep 25, 2026
@potiuk

potiuk commented Sep 25, 2026

Copy link
Copy Markdown
Member

Hello @rjgoyln - thank you for your contributions to Apache Airflow!

The Airflow community has introduced a limit of 5 open pull requests at a time for contributors without write access to the repository. You currently have 24 open pull requests, so - as a one-time step of introducing the limit - we closed the ones where maintainers have not engaged yet:

These pull requests stay open because maintainers are already engaged in them - they count towards your limit:

This is not a judgement of you or of your changes. We never told contributors before that opening many pull requests at once was a problem, so there is nothing to feel bad about - and nothing is lost: your branches, commits and the review history stay where they are.

What we ask you to do is to make your first prioritization decision: choose which of the pull requests above matter most to you, and reopen them (up to 5 open at a time, including the ones still open) with the "Reopen pull request" button or gh pr reopen <PR_NUMBER> --repo apache/airflow. Reopen the ones you are ready to follow through - keep them rebased, respond to review comments and fix failing checks.

While your pull requests are waiting for review, the most valuable thing you can do is help in other ways - reviewing other contributors' pull requests, helping with issues, and taking part in the discussions on the devlist and Slack.

Why we introduced the limit, what it means for you and how to reopen or restore a pull request is explained in https://github.com/apache/airflow/blob/main/contributing-docs/32_open_pull_request_limit.rst.


Drafted-by: Claude Code (Opus 5); reviewed by @potiuk before posting

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:scheduler closed because of open PR limit Closed as a one-time step of introducing the open pull request limit

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Scheduler bulk-fails all scheduled tasks when serialized DAG is transiently missing

2 participants