Repository navigation
Conversation
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request significantly enhances the robustness of file transfer operations by introducing a dynamic and intelligent retry mechanism. It addresses previous limitations where fixed retry parameters and cascading cancellations led to frequent failures, especially with large files over unreliable networks. The new system adapts retry behavior based on file size and isolates failures, ensuring that individual transient errors do not disrupt the entire transfer process, thereby improving overall reliability and user experience. Highlights
Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
There was a problem hiding this comment.
Code Review
This pull request introduces a robust dynamic retry mechanism with file-size-based parameters, which is a significant improvement for handling transient network errors during large file transfers. The new pkg/retrypolicy is well-designed and thoroughly tested. A key architectural improvement is replacing cascading errgroup.WithContext with independent per-goroutine retries, preventing a single failure from terminating all concurrent operations. The addition of CLI flags for retry control and progress bar updates on retry significantly enhance usability. The changes are consistently applied across all data transfer paths. I have a couple of suggestions for improvement in pkg/backend/push.go to ensure a consistent user experience during retries.
9f96098 to
1369e47
Compare
|
/gemini review |
|
@codex-code-assist review |
There was a problem hiding this comment.
Code Review
This pull request introduces a significant improvement to the application's resilience by implementing a dynamic, file-size-based retry mechanism for data transfers. The new pkg/retrypolicy is well-designed and replaces the previous fixed retry logic. A key architectural change is the move away from cascading cancellations in errgroup, which will prevent a single failure from halting all concurrent operations. The addition of CLI flags for retry control is also a valuable feature. My review focuses on a critical race condition in the progress bar handling, a minor bug in the retry backoff calculation for logging, and opportunities to reduce code duplication for better maintainability.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: eb1fc1ebc2
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| // Unknown errors default to retryable. | ||
| return true |
There was a problem hiding this comment.
Stop retrying unclassified errors as transient
This fallback makes every unexpected error retryable, and the new call sites now feed retrypolicy.Do with arbitrary local failures from paths like builder.BuildLayer and extractLayer. A permanent problem such as permission denied or no space left on device will now sit in the retry loop until the size-based deadline expires (10 minutes for small files, 60 minutes for blobs ≥10 GiB), so broken builds/pulls can appear hung for a long time instead of failing promptly.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Good point — the permanent-vs-transient split needs a deliberate allowlist (network timeouts, 5xx, connection resets, rate-limit responses, ...) rather than ad-hoc narrowing. Leaving for follow-up; I'll send a dedicated PR once we agree on the list so this one can land without scope creep.
|
/gemini review |
|
@codex-code-assist review |
There was a problem hiding this comment.
Code Review
This pull request introduces a robust and dynamic retry mechanism based on file size, significantly improving the reliability of data transfer operations. The refactoring to use a dedicated pkg/retrypolicy package centralizes retry logic, and the move to independent per-goroutine retries effectively addresses the cascading cancellation issue. The addition of CLI flags for retry configuration provides users with greater control. The comprehensive unit tests for the new retry policy are commendable, ensuring the correctness of the dynamic parameter calculations and error handling.
Overall, this is a well-designed and implemented feature that directly tackles a critical pain point for large file transfers.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: f3fbdfe46a
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
@codex-code-assist review |
1f6c0f5 to
fcdbcb0
Compare
retry-go's OnRetry callback supplies a 1-based retry attempt number, so computeBackoff(n+1, ...) logged a delay that was one doubling ahead of the actual backoff used by the retry loop. Pass n directly. Also drop the unreachable attempt == 0 branch in computeBackoff (retry-go never supplies 0) and skip that case in the test. Addresses review feedback on PR #468 (gemini). Signed-off-by: Zhao Chen <winters.zc@antgroup.com>
The 'prefer modelspec.AnnotationFilepath, fall back to the legacy dragonflyoss key' pattern was duplicated across fetch.go, fetch_by_d7y.go, and pull_by_d7y.go (six call sites total). Centralize it in backend.getAnnotationFilepath so each caller is a one-liner and future changes to the annotation resolution live in one place. Addresses review feedback on PR #468 (gemini). Signed-off-by: Zhao Chen <winters.zc@antgroup.com>
Three call sites (push, pull, processor) used 'g.Wait()' without capturing its return, so a cancelled worker's ctx.Err() could be discarded. In the edge case where all running uploads/layers succeed but queued-but-unstarted workers exit via the ctx.Done() select, errs stayed empty and the operation continued to config/manifest push (or finalized a partial descriptor set), potentially publishing an incomplete artifact. Capture the Wait() result into errs and rely on the existing ctx.Err()/len(errs) checks to return cleanly. Also change processor/base.go's select from 'return nil' to 'return ctx.Err()' so the same propagation path applies there. Additionally, align push.go's OnRetry with pull.go/fetch.go by using pb.Placeholder (designed to reset an existing bar's message and progress) instead of pb.Add for layer, config, and manifest retries. Addresses review feedback on PR #468 (codex P1, gemini). Signed-off-by: Zhao Chen <winters.zc@antgroup.com>
|
@codex-code-assist review |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: fc3ac13156
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
/gemini review |
There was a problem hiding this comment.
Code Review
This pull request introduces a specialized retrypolicy package to manage blob transfer retries, decoupling per-attempt timeouts from the retry backoff strategy. It adds CLI flags for retry configuration across the build, fetch, pull, and push commands and refactors the backend to use this new policy, improving error handling and progress tracking. A Placeholder method was added to the progress bar to handle resets during retries. Feedback suggests centralizing the progress bar reset logic within the Placeholder method and adopting the go-humanize library for byte formatting to eliminate redundant code.
There was a problem hiding this comment.
Pull request overview
This PR introduces a new pkg/retrypolicy package to provide consistent retry behavior across all blob-transfer and build paths, including dynamic per-attempt timeouts derived from file size, while also removing cascading cancellation caused by errgroup.WithContext so one failing transfer no longer cancels siblings.
Changes:
- Added
pkg/retrypolicywithDo(),IsRetryable(),ShortReason(), andComputePerAttemptTimeout()plus unit tests. - Reworked push/pull/fetch/build/processor (including Dragonfly variants) to use
retrypolicy.Doand to aggregate errors instead of canceling sibling goroutines. - Added
ProgressBar.Placeholder()to keep progress entries visible during retry backoff.
Reviewed changes
Copilot reviewed 16 out of 16 changed files in this pull request and generated 2 comments.
Show a summary per file
| File | Description |
|---|---|
| pkg/retrypolicy/retrypolicy.go | New retry policy implementation with per-attempt timeouts and retry classification. |
| pkg/retrypolicy/retrypolicy_test.go | Unit tests for timeout sizing, retryability, backoff, and Do() behavior. |
| pkg/config/push.go | Import formatting cleanup. |
| pkg/config/build.go | Import formatting cleanup. |
| pkg/backend/retry_test.go | Removed legacy retry tests tied to the old retry helper. |
| pkg/backend/push.go | Uses retrypolicy.Do and aggregates per-layer errors without cascading cancellation. |
| pkg/backend/pull.go | Uses retrypolicy.Do per layer/config/manifest and aggregates errors without canceling siblings. |
| pkg/backend/pull_by_d7y.go | Applies retrypolicy.Do to Dragonfly pull/extract path and aggregates errors. |
| pkg/backend/processor/options.go | Removes legacy defaultRetryOpts (retry is now centralized). |
| pkg/backend/processor/base.go | Uses retrypolicy.Do per-file build layer processing; aggregates errors instead of canceling. |
| pkg/backend/fetch.go | Adds retry to fetch path and aggregates errors across concurrent fetches. |
| pkg/backend/fetch_by_d7y.go | Applies retry to Dragonfly fetch/extract path and aggregates errors. |
| pkg/backend/build.go | Wraps config/manifest build steps in retrypolicy.Do. |
| pkg/backend/annotation.go | Adds getAnnotationFilepath() helper with legacy-key fallback. |
| internal/pb/pb.go | Adds ProgressBar.Placeholder() for retry/backoff UI display. |
| cmd/push.go | Flag formatting changes (multi-line flags.*Var calls). |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
0caa268 to
aba6978
Compare
aftersnow
left a comment
There was a problem hiding this comment.
Reviewed with focus on the new retry loop, concurrency behavior, and progress-bar changes. Build, go vet, and go test -race all pass locally. The core design (per-attempt timeout decoupled from a fixed retry budget) is sound and well-tested — the size-invariance test pins the key property nicely, and the msg data-race fix in internal/pb is correct.
Main findings (details inline):
- Jitter is dead code —
DelayType(BackOffDelay)drops retry-go's defaultRandomDelaycombination, soMaxJitternever applies. Concurrent layers that fail together retry in lockstep against a rate-limited registry — the exact scenario this PR targets. Empirically confirmed. - OnRetry fires after the final failed attempt — logs/progress promise a retry that never happens.
- Retry progress UX is inconsistent across push/pull/fetch paths, and
Placeholderon an aborted bar may render nothing. BeforePullLayerhook runs once-outside-retry infetch.gobut per-attempt in the other three paths.- Minor:
Configdoc drift ("CLI flags"), hand-rolledhumanizeBytes(1024 divisor with SI labels), 2024 copyright on new files, misleading// never return error to errgroupcomment.
One doc suggestion: state explicitly that the worst case per blob is ~6 × 8h ≈ 48h with no overall budget — the parent context is the only bound.
Add pkg/retrypolicy with decoupled per-attempt timeout and retry budget: - Per-attempt timeout derives from file size (10 MiB/s floor, 2x safety), clamped to [5min, 8h]; each attempt gets its own deadline. - Total attempts and per-sleep backoff cap are constants, independent of file size. - Replace cascading errgroup.WithContext cancellation with independent per-goroutine retry so one layer failure no longer cancels siblings; errors are collected and joined. - Apply retry across push, pull, fetch, build, and the Dragonfly variants. - IsRetryable uses an allowlist; unclassified errors are permanent. - Guard the progress bar message with a lock (data-race fix) and add Placeholder() for retry-backoff display. Squashed and rebased onto main; integrated with the iometrics Tracker and the updated BeforePullLayer/AfterPullLayer hook signatures. Claude-Session: https://claude.ai/code/session_01HHaMSeWpe3apQEx7n4UvRg Signed-off-by: Zhao Chen <winters.zc@antgroup.com>
Functional bugs: - retrypolicy: re-enable jitter by combining BackOffDelay with RandomDelay when jitter > 0 (guarding rand.Int63n(0) panic when it is 0). Passing BackOffDelay alone had made MaxJitter dead code, causing concurrent layers to retry in lockstep against a rate-limited registry. - retrypolicy: stop forwarding OnRetry after the final failed attempt so logs and progress bars no longer promise a retry that never happens. Consistency: - Add backend.newRetryPlaceholder and route push/pull/fetch and both Dragonfly paths through it for a uniform retry progress message. - pb.Placeholder: recreate an aborted/completed bar so the retry message actually renders, and fold in the SetRefill/EwmaSetCurrent reset. - Run BeforePullLayer/AfterPullLayer once outside the retry loop in pull, pull_by_d7y and fetch_by_d7y, matching fetch's contract. Cleanup: - humanizeBytes: delegate to go-humanize IBytes (correct KiB/MiB/GiB). - Reword Config/minThroughput docs (programmatic embedders; no CLI flags). - Bump new-file copyright to 2025; fix push.go errgroup comment. Add regression tests for the OnRetry-after-final-attempt and jitter paths. Signed-off-by: Zhao Chen <winters.zc@antgroup.com>
Signed-off-by: Zhao Chen <winters.zc@antgroup.com>
e6bbd34 to
601092c
Compare
|
Please add your SOB line in commit 601092c to pass DCO check. Otherwise lgtm. |
Signed-off-by: Zhao Chen <zhaochen.zju@gmail.com>
601092c to
c54a890
Compare
Conflicts resolved: - internal/pb/pb.go: keep main's atomic.Value message (#474) and drop the msgMu lock; Placeholder stores the message through the atomic value. - internal/pb/pb_test.go: keep main's test file and add the Placeholder tests, including the message concurrency test, on top of it. - pkg/backend/processor/base.go: keep the per-attempt retry context and add main's OnHash hook. - pkg/backend/push.go: keep the per-attempt retry context and main's pushIfNotExist signature without the prompt argument. Retry prompts use main's phase names (Pushing blob / Pushing config). Signed-off-by: Zhao Chen <zhaochen.zju@gmail.com>
PTAL. |
Summary
pkg/retrypolicypackage with decoupled per-attempt timeout and retry budget:[5 min, 8 h]). Each attempt gets its own deadline.errgroup.WithContextcancellation with independent per-goroutine retry — one layer failure no longer cancels siblings.retrypolicy.Configstruct stays available for programmatic embedders.Motivation
When pushing large model files (multi-GB to multi-TB) to OCI registries backed by rate-limited storage (e.g., Harbor + OSS),
i/o timeouterrors frequently occur. The previous retry mechanism had three problems:errgroup.WithContextmeant one timeout killed all in-flight transfers.MaxRetryTimebudget, but it covered both in-flight transfer time and inter-attempt sleeps. With a wall-clock that scales with file size, a slow first attempt could consume the whole budget, leaving no room for retries — exactly when retries matter most.The current design splits the timing concerns into two independent constants, with no per-invocation overrides. We considered exposing them as CLI flags (
--retry-attempts,--per-attempt-timeout) but found they were operational settings that rarely vary per invocation; YAGNI says skip until a real user case shows up.Design
ComputePerAttemptTimeout(fileSize)—file_size / 10 MiB/s × 2, clamped to[5 min, 8 h]DefaultMaxAttempts = 6DefaultMaxBackoff = 2 minDefaultInitialDelay = 5 sPer-attempt deadlines are derived inside
Do()viacontext.WithTimeout(ctx, perAttemptTimeout); the parentctxis reserved for user cancellation. ADeadlineExceedederror under a live parent context is reclassified as retryable, so a single transfer timeout no longer short-circuits the retry loop.Examples (
ComputePerAttemptTimeout):Changes
pkg/retrypolicy/Do(),IsRetryable(),ShortReason(),ComputePerAttemptTimeout()pkg/backend/push.goretrypolicy.Dofor config/manifestpkg/backend/pull.gopb.Placeholder()for retry progress displaypkg/backend/fetch.gopkg/backend/build.gopkg/backend/processor/base.gopkg/backend/pull_by_d7y.go,fetch_by_d7y.gointernal/pb/pb.goPlaceholder()method for retry backoff displaypkg/backend/retry.godefaultRetryOpts)No new CLI flags or config fields.
Test plan
pkg/retrypolicy/retrypolicy_test.go, including a size-invariance test that pins the design's core property: total retry wall-clock does not depend on file size.go vet ./...clean.go test -race ./pkg/retrypolicy/...clean.make lintclean (golangci-lint v2.5.0).