Skip to content

Time-box and retry pre-extras SDK downloads - #71544

Merged
shahar1 merged 1 commit into
apache:mainfrom
potiuk:retry-pre-extras-downloads
Oct 3, 2026
Merged

shahar1 merged 1 commit into
apache:mainfrom
potiuk:retry-pre-extras-downloads

Conversation

@potiuk

@potiuk potiuk commented Aug 13, 2026

Copy link
Copy Markdown
Member

The lowest-dependency providers job fails when IBM's download server stops answering: the ibm.mq pre-extras manifest fetch passed no socket timeout, so each of the five routes (DNS plus the four published anycast IPs) sat in the kernel's TCP connect timeout, and the single pass over them had no retry. One transient upstream outage took the whole job down — see the Errno 110 failures in this run.

Every attempt now uses a 20s socket timeout, and the full set of routes is retried in 3 rounds with a short sleep between them, so a stalled transfer or a blackholed IP recovers in seconds. A checksum mismatch is still fatal on the first attempt — that means the manifest disagrees with what upstream serves, which retrying cannot fix.


Was generative AI tooling used to co-author this PR?
  • Yes — Claude Code (Opus 5)

Generated-by: Claude Code (Opus 5) following the guidelines

A lowest-dependency providers job failed because IBM's download server
stopped answering: with no socket timeout, each of the five routes (DNS
plus the four published anycast IPs) sat in the kernel's TCP connect
timeout, and the single pass over them had no retry, so a transient
upstream outage took the whole job down after minutes of waiting.

Third-party servers hosting these SDKs are the least reliable part of
the run, and an outage measured in seconds should not cost a CI job.
@potiuk

potiuk commented Aug 17, 2026 •

Copy link
Copy Markdown
Member Author

Green and self-contained: the pre-extras SDK download now has a time box and retries instead of being able to hang the build.

@dabla you have both worked on this script — could one of you take a look?


Drafted-by: Claude Code (Opus 5); reviewed by @potiuk before posting

@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

This pull request has been automatically marked as stale because it has not had recent activity. It will be closed in 5 days if no further activity occurs. Thank you for your contributions.

@github-actions github-actions Bot added the stale Stale PRs per the .github/workflows/stale.yml policy file label Oct 2, 2026
@shahar1 shahar1 removed the stale Stale PRs per the .github/workflows/stale.yml policy file label Oct 3, 2026
@shahar1
shahar1 merged commit da53196 into apache:main Oct 3, 2026
155 checks passed
@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

Backport successfully created: v3-3-test

Note: As of Merging PRs targeted for Airflow 3.X
the committer who merges the PR is responsible for backporting the PRs that are bug fixes (generally speaking) to the maintenance branches.

In matter of doubt please ask in #release-management Slack channel.

Status Branch Result
✅ v3-3-test PR Link

github-actions Bot pushed a commit to aws-mwaa/upstream-to-airflow that referenced this pull request Oct 3, 2026
A lowest-dependency providers job failed because IBM's download server
stopped answering: with no socket timeout, each of the five routes (DNS
plus the four published anycast IPs) sat in the kernel's TCP connect
timeout, and the single pass over them had no retry, so a transient
upstream outage took the whole job down after minutes of waiting.

Third-party servers hosting these SDKs are the least reliable part of
the run, and an outage measured in seconds should not cost a CI job.
(cherry picked from commit da53196)

Co-authored-by: Jarek Potiuk <jarek@potiuk.com>
aws-airflow-bot pushed a commit to aws-mwaa/upstream-to-airflow that referenced this pull request Oct 3, 2026
A lowest-dependency providers job failed because IBM's download server
stopped answering: with no socket timeout, each of the five routes (DNS
plus the four published anycast IPs) sat in the kernel's TCP connect
timeout, and the single pass over them had no retry, so a transient
upstream outage took the whole job down after minutes of waiting.

Third-party servers hosting these SDKs are the least reliable part of
the run, and an outage measured in seconds should not cost a CI job.
(cherry picked from commit da53196)

Co-authored-by: Jarek Potiuk <jarek@potiuk.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants