Skip to content

Deferrable Dataproc triggers swallow CancelledError during triggerer migration, mass-failing deferred tasks (same as #63730 for BigQuery) #74087

Description

@krisztiansala

Apache Airflow version

3.2.2 (also present on main / google provider 22.6.0)

What happened

When a running triggerer job's heartbeat goes stale past [triggerer] triggerer_health_check_threshold (30s default) — e.g. because the TriggerRunner subprocess event loop is temporarily blocked on a busy/undersized triggerer, or because the new runner_health_check_threshold watchdog deliberately skips heartbeats when the runner is silent — a sibling triggerer's assign_unassigned() steals all of its trigger rows. The original triggerer, still alive, then diffs load_triggers(), sees the rows reassigned away, and calls task.cancel() on every one of its in-flight trigger coroutines.

DataprocSubmitTrigger.run() (and DataprocBatchTrigger, DataprocClusterTrigger — same pattern) catches the asyncio.CancelledError, awaits safe_to_cancel() (supervisor comms call via greenback), finds the TI still DEFERRED, skips cancel_job — and then does not re-raise. The coroutine exits cleanly having emitted zero events.

cleanup_finished_triggers() only recognizes a cancellation when task.result() raises CancelledError. A swallowed CancelledError makes the exit look like a crash:

Trigger exited without sending an event. Dependent tasks will be failed.

→ Trigger.submit_failure() → every dependent deferred task is rescheduled with next_method=__fail__ and fails on the worker with TaskDeferralError: Trigger failure, while the Dataproc job keeps running to completion.

So a transient triggerer stall is amplified into a mass failure of all deferred tasks it hosted — the exact opposite of what trigger migration is supposed to achieve (transparent failover).

This is the same bug shape fixed for BigQueryInsertJobTrigger in #63730 ("Fix BigQueryInsertJobTrigger not propagating CancelledError"): the except CancelledError block must re-raise after cleanup so the framework can tell "cancelled during migration" from "crashed".

How to reproduce

  1. Two triggerers, several DataprocSubmitJobOperator(deferrable=True) tasks deferred.
  2. Stall one triggerer job's heartbeat > triggerer_health_check_threshold (e.g. suspend the process, or saturate the runner event loop so runner_health_check_threshold trips heartbeat suppression).
  3. Sibling steals the trigger rows; original triggerer cancels its coroutines.
  4. All affected tasks fail with TaskDeferralError: Trigger failure instead of transparently migrating.

Observed in production at scale on Cloud Composer 3 (composer-3-airflow-3.2.2-build.2): bursts of 12+ and 3+ deferred tasks failing within a second, triggerer log showing 12x Trigger exited without sending an event + Got response for unknown request frame warnings + Task cancelling ... coro=<greenback_shim()> entries immediately beforehand.

Suggested fix

In providers/google/src/airflow/providers/google/cloud/triggers/dataproc.py, add a bare raise at the end of the except asyncio.CancelledError: blocks (all three trigger classes), so cancellation propagates and cleanup_finished_triggers takes the expected except (CancelledError, ...) -> del + continue path — no submit_failure, sibling triggerer's copy resumes the task normally.

Possibly also worth auditing other google provider triggers (BigQuery excepted — fixed in #63730) and other providers for the same swallowed-CancelledError pattern.

Activity

  1. boring-cyborg commented on Oct 2, 2026

    @boring-cyborg

    Thanks for opening your first issue here! Be sure to follow the issue template! If you are willing to raise PR to address this issue please do so, no need to wait for approval.

  2. kadubhumika commented on Oct 2, 2026

    @kadubhumika
    Contributor

    @krisztiansala Hi ! I would love to take on this issue! Its interesting and backend oriented As noted in the description, this follows the same pattern as #63730 for BigQuery. I will inspect the code in providers/google/src/airflow/providers/google/cloud/triggers/dataproc.py, ensure asyncio.CancelledError is properly re-raised across the three Dataproc triggers, and open a PR shortly.

  3. krisztiansala commented on Oct 2, 2026

    @krisztiansala
    Author

    thank you @kadubhumika

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    kind:bugThis is a clearly a bugprovider:googleGoogle (including GCP) related issues

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions