Repository navigation
Conversation
077e4a0 to
d654d6e
Compare
The scheduler's hand-rolled DB-retry loops are meant to ride out a deadlock and carry on. They cannot: Session._flush rolls its subtransaction back on any exception, so a failed flush leaves the session deactivated and the next attempt's first statement raises PendingRollbackError, which the retry predicate does not match and which therefore escapes the loop. The timer driving these sweeps does not swallow exceptions, so one transient deadlock takes the scheduler down instead of costing a single tick. retry_db_transaction has rolled back for exactly this reason since it was written. The awaiting-input sweep never did, and the orphaned-task sweep covered only one of the error types its loop retries.
d654d6e to
a60617a
Compare
|
Hello @rjgoyln - thank you for your contributions to Apache Airflow! The Airflow community has introduced a limit of 5 open pull requests at a time for contributors without write access to the repository. You currently have 24 open pull requests, so - as a one-time step of introducing the limit - we closed the ones where maintainers have not engaged yet:
These pull requests stay open because maintainers are already engaged in them - they count towards your limit:
This is not a judgement of you or of your changes. We never told contributors before that opening many pull requests at once was a problem, so there is nothing to feel bad about - and nothing is lost: your branches, commits and the review history stay where they are. What we ask you to do is to make your first prioritization decision: choose which of the pull requests above matter most to you, and reopen them (up to 5 open at a time, including the ones still open) with the "Reopen pull request" button or While your pull requests are waiting for review, the most valuable thing you can do is help in other ways - reviewing other contributors' pull requests, helping with issues, and taking part in the discussions on the devlist and Slack. Why we introduced the limit, what it means for you and how to reopen or restore a pull request is explained in https://github.com/apache/airflow/blob/main/contributing-docs/32_open_pull_request_limit.rst. Drafted-by: Claude Code (Opus 5); reviewed by @potiuk before posting |
Summary
check_awaiting_input_timeoutsandadopt_or_reset_orphaned_tasksboth run insiderun_with_db_retries, so a deadlock should cost a scheduler tick rather than the scheduler. Neither loop survives one.A failed flush deactivates the session:
Session._flushrolls its subtransaction back on any exception, so a plainStaleDataErrorpoisons it as surely as a DBAPI error. The retry's first statement then raisesPendingRollbackError, which the predicate does not match, so it escapes the loop — and the awaiting-input sweep sits on a timer withoutnon_fatal=True, taking the scheduler loop with it.retry_db_transactionrolls back for exactly this reason; neither hand-rolled loop picked it up.Change
check_awaiting_input_timeouts, asretry_db_transactiondoes.(DBAPIError, StaleDataError).Both guards now name exactly the tuple
run_with_db_retriesretries. That widens no retries — tenacity already retried everyDBAPIErrorin the orphaned-task sweep; those retries just died on the next statement.The sweep body is unchanged apart from the indent under the new
try:;check_trigger_timeoutsis the third loop of this shape, and #72595 covers it.Tests
Both tests drive a session that deactivates on a failed flush the way SQLAlchemy does; all four cases fail on
mainwithPendingRollbackError.Was generative AI tooling used to co-author this PR?
Generated-by: Claude Code (Opus 5) following the guidelines