Skip to content

Make the scheduler's HITL and orphan sweeps retry transient DB errors - #72596

Closed
rjgoyln wants to merge 1 commit into
apache:mainfrom
rjgoyln:fix/hitl-timeout-sweep-rollback
Closed

rjgoyln wants to merge 1 commit into
apache:mainfrom
rjgoyln:fix/hitl-timeout-sweep-rollback

Conversation

@rjgoyln

@rjgoyln rjgoyln commented Sep 6, 2026 •

Copy link
Copy Markdown
Contributor

Summary

check_awaiting_input_timeouts and adopt_or_reset_orphaned_tasks both run inside run_with_db_retries, so a deadlock should cost a scheduler tick rather than the scheduler. Neither loop survives one.

A failed flush deactivates the session: Session._flush rolls its subtransaction back on any exception, so a plain StaleDataError poisons it as surely as a DBAPI error. The retry's first statement then raises PendingRollbackError, which the predicate does not match, so it escapes the loop — and the awaiting-input sweep sits on a timer without non_fatal=True, taking the scheduler loop with it.

retry_db_transaction rolls back for exactly this reason; neither hand-rolled loop picked it up.

Change

  • Roll back before re-raising in check_awaiting_input_timeouts, as retry_db_transaction does.
  • Widen the orphaned-task sweep's guard to (DBAPIError, StaleDataError).

Both guards now name exactly the tuple run_with_db_retries retries. That widens no retries — tenacity already retried every DBAPIError in the orphaned-task sweep; those retries just died on the next statement.

The sweep body is unchanged apart from the indent under the new try:; check_trigger_timeouts is the third loop of this shape, and #72595 covers it.

Tests

Both tests drive a session that deactivates on a failed flush the way SQLAlchemy does; all four cases fail on main with PendingRollbackError.


Was generative AI tooling used to co-author this PR?
  • Yes — Claude Code (Opus 5)

Generated-by: Claude Code (Opus 5) following the guidelines

@rjgoyln
rjgoyln force-pushed the fix/hitl-timeout-sweep-rollback branch 4 times, most recently from 077e4a0 to d654d6e Compare September 9, 2026 14:32
The scheduler's hand-rolled DB-retry loops are meant to ride out a deadlock
and carry on. They cannot: Session._flush rolls its subtransaction back on
any exception, so a failed flush leaves the session deactivated and the next
attempt's first statement raises PendingRollbackError, which the retry
predicate does not match and which therefore escapes the loop. The timer
driving these sweeps does not swallow exceptions, so one transient deadlock
takes the scheduler down instead of costing a single tick.

retry_db_transaction has rolled back for exactly this reason since it was
written. The awaiting-input sweep never did, and the orphaned-task sweep
covered only one of the error types its loop retries.
@potiuk

potiuk commented Sep 25, 2026

Copy link
Copy Markdown
Member

Hello @rjgoyln - thank you for your contributions to Apache Airflow!

The Airflow community has introduced a limit of 5 open pull requests at a time for contributors without write access to the repository. You currently have 24 open pull requests, so - as a one-time step of introducing the limit - we closed the ones where maintainers have not engaged yet:

These pull requests stay open because maintainers are already engaged in them - they count towards your limit:

This is not a judgement of you or of your changes. We never told contributors before that opening many pull requests at once was a problem, so there is nothing to feel bad about - and nothing is lost: your branches, commits and the review history stay where they are.

What we ask you to do is to make your first prioritization decision: choose which of the pull requests above matter most to you, and reopen them (up to 5 open at a time, including the ones still open) with the "Reopen pull request" button or gh pr reopen <PR_NUMBER> --repo apache/airflow. Reopen the ones you are ready to follow through - keep them rebased, respond to review comments and fix failing checks.

While your pull requests are waiting for review, the most valuable thing you can do is help in other ways - reviewing other contributors' pull requests, helping with issues, and taking part in the discussions on the devlist and Slack.

Why we introduced the limit, what it means for you and how to reopen or restore a pull request is explained in https://github.com/apache/airflow/blob/main/contributing-docs/32_open_pull_request_limit.rst.


Drafted-by: Claude Code (Opus 5); reviewed by @potiuk before posting

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:scheduler closed because of open PR limit Closed as a one-time step of introducing the open pull request limit

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants