test(diffusion): HeavyTimeout-tag 5 verified compute-bound models - #1749
Conversation
/#1305) These diffusion models exceed the 120s [Fact(Timeout)] training probe IN ISOLATION (verified solo, one model per process), so they're genuinely compute-bound — not parallel-starvation flakes. Tag them [Trait("Category","HeavyTimeout")] so the default PR shard gate (filtered &Category!=HeavyTimeout) skips them and the nightly HeavyTimeout lane runs them instead: PixArtDeltaLCM, PlaygroundV3, ARDiffusion, MoMask, SiDDiT Deliberately NOT tagged (verified to PASS solo in ~10-15s — they only timed out under parallel core-starvation / WeightRegistry OOM accumulation, which is the AiDotNet.Tensors #714 foundation-memory bug, not a per-model compute cost): DreamFusion, InstructVid2Vid, CogVideo, VideoCrafter2, FateZero, FlowVid Verified the exclusion: `FullyQualifiedName~PixArtDeltaLCMModelTests&Category!=HeavyTimeout` matches 0 tests after tagging. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
|
Warning Review limit reachedYou’ve reached a temporary PR review limit under our Fair Usage Limits Policy. Next review available in: 42 minutes Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Pro Run ID: 📒 Files selected for processing (5)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
…) via HeavyTimeout (#1762) * test(ci): HeavyTimeout-tag foundation-scale diffusion models in 2 OOM-cancelled shards Both shards were SIGTERM-cancelled (exit 143, 'runner received a shutdown signal') on the 16 GB github-hosted runner, not by assertion failures — the classic 'Cancelled-runner shards (Diffusion)' foundation-scale OOM pattern. The test base already disposes each model + GC-reclaims between tests, so this is a single-model peak, not a leak. ModelFamily - Diffusion SA-SD: 8 model classes instantiate at full default (paper) scale via default ctors (SANA, SANASprint, SASTD, SCott, SD3Flash, SD3Inpainting, SD3Turbo, SDXLLightning) — SD3/SDXL-based, billions of params. Tag [Category=HeavyTimeout] (+ FoundationScaleSerial) so the default gate (Category!=HeavyTimeout) skips them and they run in the nightly lane, matching the already-tagged SDXLTurbo/SDXLInpainting siblings. Unit - 03b Control/DDPM/Preprocessor: ControlModelContractTests constructs each model and asserts ParameterCount>0. The Flux (~12B ~= 48 GB fp32) and SD3 backbones cannot fit 16 GB even to construct. Per-method HeavyTimeout on the 4 Flux/SD3 methods (ControlNetFlux ctor + Clone, ControlNetSD3, ControlNetPlusPlusFlux); the SD-scale ControlNets and the lightweight guidance tests stay on the PR gate. Same established fix as #1744/#1749/#1706; models still run fully in the nightly HeavyTimeout lane. Build verified (net10.0). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(ci): align SANA HeavyTimeout tag with #1744 to avoid a merge conflict #1744 also tags SANAModelTests.cs (with different comment text). Make this PR's SANA edit byte-identical to #1744's so the two PRs 3-way-merge cleanly instead of conflicting. The other 7 SA-SD models + the 03b ControlModelContractTests Flux/SD3 methods remain #1762-exclusive gaps that #1744 does not cover. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
What
Tags 5 diffusion model test classes
[Trait("Category","HeavyTimeout")]so the default PR shard gate (which filters&Category!=HeavyTimeout) skips them and the nightly HeavyTimeout lane runs them instead — unblocking the diffusion shards from these models' 120s training timeouts.Tagged (verified compute-bound — exceed the 120s
[Fact(Timeout)]training probe IN ISOLATION, one model per process):PixArtDeltaLCM,PlaygroundV3,ARDiffusion,MoMask,SiDDiTDiscipline: what was deliberately NOT tagged
Several models that timed out in the parallel family run pass solo in ~10–15s — they're not compute-bound, they were starved/OOM'd by the concurrent run. Tagging them would wrongly hide passing tests, so they're left alone:
DreamFusion, InstructVid2Vid, CogVideo, VideoCrafter2, FateZero, FlowVidTheir parallel-only failure is the AiDotNet.Tensors #714 foundation-memory bug (WeightRegistry streaming-pool accumulation → ~18 GB OOM-hang), not a per-model cost — fixing #714 is what greens those.
Verification
FullyQualifiedName~PixArtDeltaLCMModelTests&Category!=HeavyTimeout→ 0 tests matched after tagging.Scope
First pass over the compute-bound diffusion models surfaced so far. Exhaustive enumeration of every heavy diffusion model is itself blocked by #714 (the full family can't complete a parallel run to enumerate them), so further tags will follow as CI surfaces them. Part of the #1706 / #1305 shard-greening effort.
🤖 Generated with Claude Code