Repository navigation
Fix Databricks hook dropping tasks beyond the first page of a run - #72304
Merged
Merged
Conversation
The Jobs API 2.2 returns at most 100 entries of a run's tasks and job_clusters per page and hands back a token for the rest, but get_run read only the first page. Callers that inspect the task list, such as the failed-task error extraction behind a task's failure message, therefore saw an arbitrary subset of it: the entries do not come back in the order they were declared, so which ones went missing was not predictable.
1 task done
Contributor
Author
|
@eladkal could you take a look when you have a moment, or point it at whoever is better placed?
Drafted-by: Claude Code (Opus 5); reviewed by @moomindani before posting |
eladkal
approved these changes
Sep 9, 2026
imrichardwu
pushed a commit
to imrichardwu/airflow
that referenced
this pull request
Sep 11, 2026
…ache#72304) The Jobs API 2.2 returns at most 100 entries of a run's tasks and job_clusters per page and hands back a token for the rest, but get_run read only the first page. Callers that inspect the task list, such as the failed-task error extraction behind a task's failure message, therefore saw an arbitrary subset of it: the entries do not come back in the order they were declared, so which ones went missing was not predictable.
xvega
pushed a commit
to xvega/airflow
that referenced
this pull request
Sep 13, 2026
…ache#72304) The Jobs API 2.2 returns at most 100 entries of a run's tasks and job_clusters per page and hands back a token for the rest, but get_run read only the first page. Callers that inspect the task list, such as the failed-task error extraction behind a task's failure message, therefore saw an arbitrary subset of it: the entries do not come back in the order they were declared, so which ones went missing was not predictable.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
DatabricksHook.get_runand its async twin read only the first page of2.2/jobs/runs/get. The Jobs API returns at most 100 entries of a run'stasksandjob_clustersper page and hands back anext_page_tokenfor the rest, so on a job with more tasks than that the run came back carrying an arbitrary subset of them — the entries are not returned in declaration order, so which ones went missing was not predictable.The visible effect is in
extract_failed_task_errors[_async], which walksrun_info["tasks"]to attach each failed task's own error to the message the task fails with: a failure on a later page falls back to the generic run state message instead.get_run_tasksalready paginated, so it now delegates toget_runand the two cannot drift apart again.Verified against a live workspace, using a job of 101
condition_taskentries (no compute) plus one failing notebook task:get_run()["tasks"]returned 100 entries whileget_run_tasks()returned 103, and the keys missing from the former were an arbitrary triojob_clustersis paginated too rather than repeated on each page — a job with two job clusters returns both on page 1 and none on page 2 — so merging it is correctBoth new unit tests fail without the change (
call_count 1 == 2), and the pre-existingget_run_tasksmulti-page test passes unchanged.Was generative AI tooling used to co-author this PR?
Generated-by: Claude Code (Opus 5) following the guidelines