Repository navigation
Time-box and retry pre-extras SDK downloads - #71544
Merged
Merged
Conversation
A lowest-dependency providers job failed because IBM's download server stopped answering: with no socket timeout, each of the five routes (DNS plus the four published anycast IPs) sat in the kernel's TCP connect timeout, and the single pass over them had no retry, so a transient upstream outage took the whole job down after minutes of waiting. Third-party servers hosting these SDKs are the least reliable part of the run, and an outage measured in seconds should not cost a CI job.
potiuk
requested review from
amoghrajesh,
ashb,
bugraoz93,
gopidesupavan,
jason810496 and
jscheffl
as code owners
August 13, 2026 10:28
Member
Author
1 task done
1 task done
Contributor
|
This pull request has been automatically marked as stale because it has not had recent activity. It will be closed in 5 days if no further activity occurs. Thank you for your contributions. |
1 task done
shahar1
approved these changes
Oct 3, 2026
Contributor
Backport successfully created: v3-3-testNote: As of Merging PRs targeted for Airflow 3.X In matter of doubt please ask in #release-management Slack channel.
|
github-actions Bot
pushed a commit
to aws-mwaa/upstream-to-airflow
that referenced
this pull request
Oct 3, 2026
A lowest-dependency providers job failed because IBM's download server stopped answering: with no socket timeout, each of the five routes (DNS plus the four published anycast IPs) sat in the kernel's TCP connect timeout, and the single pass over them had no retry, so a transient upstream outage took the whole job down after minutes of waiting. Third-party servers hosting these SDKs are the least reliable part of the run, and an outage measured in seconds should not cost a CI job. (cherry picked from commit da53196) Co-authored-by: Jarek Potiuk <jarek@potiuk.com>
aws-airflow-bot
pushed a commit
to aws-mwaa/upstream-to-airflow
that referenced
this pull request
Oct 3, 2026
A lowest-dependency providers job failed because IBM's download server stopped answering: with no socket timeout, each of the five routes (DNS plus the four published anycast IPs) sat in the kernel's TCP connect timeout, and the single pass over them had no retry, so a transient upstream outage took the whole job down after minutes of waiting. Third-party servers hosting these SDKs are the least reliable part of the run, and an outage measured in seconds should not cost a CI job. (cherry picked from commit da53196) Co-authored-by: Jarek Potiuk <jarek@potiuk.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The lowest-dependency providers job fails when IBM's download server stops answering: the
ibm.mqpre-extras manifest fetch passed no socket timeout, so each of the five routes (DNS plus the four published anycast IPs) sat in the kernel's TCP connect timeout, and the single pass over them had no retry. One transient upstream outage took the whole job down — see theErrno 110failures in this run.Every attempt now uses a 20s socket timeout, and the full set of routes is retried in 3 rounds with a short sleep between them, so a stalled transfer or a blackholed IP recovers in seconds. A checksum mismatch is still fatal on the first attempt — that means the manifest disagrees with what upstream serves, which retrying cannot fix.
Was generative AI tooling used to co-author this PR?
Generated-by: Claude Code (Opus 5) following the guidelines