Repository navigation
Scheduler bulk-fails all scheduled tasks when serialized DAG is transiently missing #62050
Description
Activity
- addedkind:bugThis is a clearly a bugThis is a clearly a bugneeds-triagelabel for new issues that we didn't triage yetlabel for new issues that we didn't triage yet
on Feb 16, 2026 Thanks for opening your first issue here! Be sure to follow the issue template! If you are willing to raise PR to address this issue please do so, no need to wait for approval.
- added 9 commits that reference this issue
on Mar 4, 2026 24 remaining items
- added 2 commits that reference this issue
on Apr 13, 2026 - addedaffected_version:3.2Use for reporting issues with 3.2Use for reporting issues with 3.2and removedneeds-triagelabel for new issues that we didn't triage yetlabel for new issues that we didn't triage yet
on May 11, 2026 Hi @kaxil, @jscheffl, and @potiuk — regarding the recent closure of #72243 and the concerns raised in #62878, I did a deep dive and reproduced this locally on Postgres 14 against
mainwithout mocking. The results actually prove your instincts right, but change the shape of the fix. I'd like to align on the direction before opening a PR.1. The current
UPDATEis destructive, not a safety netWhen a serialized DAG is missing, the current bulk
UPDATEcauses:- Cross-run blast radius: It wrongly fails SCHEDULED tasks across all runs of the DAG, not just the affected one.
-
Silent UI failures: Tasks go straight to FAILED. It bypasses
handle_failureentirely—no retries, no callbacks, no logs. -
$O(N)$ DB spam: Every TI of the missing DAG triggers its own full-tableUPDATEin the batch. -
Stuck runs: The tasks are destroyed, but the DagRun remains stuck in
RUNNINGanyway.
2. The Starvation Trap: Why "pure skip" (#62878 / #72243) is worse
The current
UPDATEaccidentally acts as a starvation filter. If we just remove it ("pure skip"), the broken tasks loop forever, locking out healthy DAGs.Here is a 2-pass scheduler test with a broken DAG (
max_tis=2) and a healthy DAG:Scenario PASS 1 queued PASS 2 queued Broken DAG's tasks Healthy DAG's task Current main[][healthy_task]FAILED (Wrongly killed) Queued after 1 loop delay Pure Skip [][]SCHEDULED Never queued (Cluster Starvation!) Skip + starved_dags[healthy_task][]SCHEDULED Queued immediately 3. The Proposed Fix
When
get_dag_for_run()returnsNonein_task_concurrency_allows_execution:- Leave TIs as
SCHEDULED(stop the destructive UPDATE). - Add
dag_idtostarved_dagsto prevent the cluster starvation shown above. - Track the missing DAG locally to log the error only once per DAG per batch, not once per TI.
4. Scope & Next Steps
The remaining concern (from #62878) is that a permanently deleted DAG leaves TIs SCHEDULED forever. Since this is a DagRun-level gap already being addressed by #70056, I plan to scope my PR strictly to the concurrency and starvation issues above.
Does this plan look good to you? I have the local reproduction ready as a test and can open the PR if you agree with this direction.
- added 2 commits that reference this issue
on Sep 9, 2026 - added 7 commits that reference this issue
on Sep 13, 2026
Apache Airflow version
Other Airflow 3 version (please specify below)
If "Other Airflow 3 version" selected, which one?
3.0.6, still on main
What happened?
My instance has been hitting intermittent task failures on MWAA (Airflow 3.0.6, ~150 DAGs, 4 schedulers). Tasks failed in bulk with no obvious cause but succeeded on manual retry. I noticed this on
scheduler_job_runner.py:When the scheduler can't find a DAG in the
serialized_dagtable, it does this:It sets every SCHEDULED task instance for that DAG to FAILED.
With PR #58259 and #56422, it probably happens less often but the bulk-failure issue has never been addressed
What you think should happen instead?
The scheduler could skip scheduling that DAG for the current iteration and try again next time, instead of immediately failing everything.
I thought of logging a warning instead of error, tracking a counter per DAG, and only failing tasks after several consecutive misses to distinguish transient gaps from genuinely missing DAGs. What do you think about this solution?
PR #55126 tried something similar for stale DAGs (skipping and continuing)
How to reproduce
Operating System
AWS MWAA
Versions of Apache Airflow Providers
n/a
Deployment
Amazon (AWS) MWAA
Deployment details
Anything else?
would appreciate guidance on the preferred approach (retry counter or just ignoring)
Are you willing to submit PR?
Code of Conduct