Skip to content

Fix Databricks hook dropping tasks beyond the first page of a run - #72304

Merged
eladkal merged 1 commit into
apache:mainfrom
moomindani:fix-databricks-get-run-pagination
Sep 9, 2026
Merged

eladkal merged 1 commit into
apache:mainfrom
moomindani:fix-databricks-get-run-pagination

Conversation

@moomindani

Copy link
Copy Markdown
Contributor

DatabricksHook.get_run and its async twin read only the first page of 2.2/jobs/runs/get. The Jobs API returns at most 100 entries of a run's tasks and job_clusters per page and hands back a next_page_token for the rest, so on a job with more tasks than that the run came back carrying an arbitrary subset of them — the entries are not returned in declaration order, so which ones went missing was not predictable.

The visible effect is in extract_failed_task_errors[_async], which walks run_info["tasks"] to attach each failed task's own error to the message the task fails with: a failure on a later page falls back to the generic run state message instead.

get_run_tasks already paginated, so it now delegates to get_run and the two cannot drift apart again.

Verified against a live workspace, using a job of 101 condition_task entries (no compute) plus one failing notebook task:

  • before: get_run()["tasks"] returned 100 entries while get_run_tasks() returned 103, and the keys missing from the former were an arbitrary trio
  • after: both return the complete list, with the run-level fields unchanged
  • job_clusters is paginated too rather than repeated on each page — a job with two job clusters returns both on page 1 and none on page 2 — so merging it is correct

Both new unit tests fail without the change (call_count 1 == 2), and the pre-existing get_run_tasks multi-page test passes unchanged.


Was generative AI tooling used to co-author this PR?
  • Yes — Claude Code (Opus 5)

Generated-by: Claude Code (Opus 5) following the guidelines

The Jobs API 2.2 returns at most 100 entries of a run's tasks and
job_clusters per page and hands back a token for the rest, but get_run
read only the first page. Callers that inspect the task list, such as the
failed-task error extraction behind a task's failure message, therefore
saw an arbitrary subset of it: the entries do not come back in the order
they were declared, so which ones went missing was not predictable.
@moomindani

Copy link
Copy Markdown
Contributor Author

@eladkal could you take a look when you have a moment, or point it at whoever is better placed?

get_run reads only the first page of a run, so on a job with more than 100 tasks an arbitrary subset goes missing and a failed task's own error falls back to the generic run message. Green, and verified against a live workspace.


Drafted-by: Claude Code (Opus 5); reviewed by @moomindani before posting

@eladkal
eladkal merged commit 069a973 into apache:main Sep 9, 2026
83 checks passed
imrichardwu pushed a commit to imrichardwu/airflow that referenced this pull request Sep 11, 2026
…ache#72304)

The Jobs API 2.2 returns at most 100 entries of a run's tasks and
job_clusters per page and hands back a token for the rest, but get_run
read only the first page. Callers that inspect the task list, such as the
failed-task error extraction behind a task's failure message, therefore
saw an arbitrary subset of it: the entries do not come back in the order
they were declared, so which ones went missing was not predictable.
xvega pushed a commit to xvega/airflow that referenced this pull request Sep 13, 2026
…ache#72304)

The Jobs API 2.2 returns at most 100 entries of a run's tasks and
job_clusters per page and hands back a token for the rest, but get_run
read only the first page. Callers that inspect the task list, such as the
failed-task error extraction behind a task's failure message, therefore
saw an arbitrary subset of it: the entries do not come back in the order
they were declared, so which ones went missing was not predictable.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants