Skip to content

ci: split OOMing model shards finer + cover ~780 never-run tests (Document AI, Finance, integration suite) - #1540

Merged
ooples merged 1 commit into
masterfrom
ci/shard-split-oom-coverage
Jun 8, 2026
Merged

ooples merged 1 commit into
masterfrom
ci/shard-split-oom-coverage

Conversation

@ooples

@ooples ooples commented Jun 8, 2026 •

Copy link
Copy Markdown
Owner

Summary

Fixes two problems in the test CI matrix (sonarcloud.yml):

1. OOM cancellations on the heavy ModelFamily shards

The 16 GB ubuntu-latest runner is killed ("received a shutdown signal" 2–6 min in, 0 test failures) on the Diffusion / Generated / NeuralNetworks shards. Each instantiates dozens of paper-scale models (~2.6 GB each: 880 MB weights + 1.76 GB Adam state). The existing mitigations — per-test GC.Collect (already in NeuralNetworkModelTestBase.DisposeAsync) and serialize+Workstation-GC for heavy shards — weren't enough because the cumulative distinct-model footprint per shard was too large.

Fix: split finer so each shard loads fewer model classes, and serialize them:

  • Diffusion 3 → 6 shards (A-C / D-I / J-M / N-R / S / T-Z; S alone = 49 models, the heaviest letter)
  • NeuralNetworks 1 → 2 (A-L / M-Z, 85 models)
  • Generated Layers 1 → 2 (A-M / N-Z, ~1670 methods)

2. Entire test trees ran in ZERO shards (the "missing models" you noticed)

No shard matched these, so they never executed in CI:

  • AiDotNet.Tests.IntegrationTests (~686 files) — incl. Document AI (IntegrationTests/Document/*: LayoutAware, PixelToSequence, VisionLanguage, GraphBased), Finance, ComputerVision, Audio, Video, MetaLearning, Optimizers, Preprocessing, Statistics, … → added 10 gap-free letter-grouped Integration shards (A-B, C, D, E-G, H-L, M, N-O, P-Q, R, S, T-Z — partitions all of A–Z).
  • ModelFamily CodeModel / Forecasting / Segmentation / Survival (CodeBERT, MGTSD, SwinUNETR, RandomSurvivalForest) → one combined shard.
  • Top-level FederatedLearning (50 files), Audio, Onnx, Tokenization, ComponentTests, AdversarialRobustness, EndToEndTests → 3 shards.

Model-instantiating Integration/TopLevel groups are serialized via $heavyShards; light ones run parallel.

Heads-up

The newly-covered shards will likely surface pre-existing failures in model families that were never run in CI — that's the intent (it's why Document AI / Financial AI looked "missing"). Those become tracked follow-ups; this PR establishes the coverage and stops the OOM cancellations.

Verification

  • Workflow YAML parses; the push-step bash is unchanged in structure.
  • Coverage is gap-free by construction (letter partitions span A–Z); ci-shard-closure-policy.yml is an issue-close guard and needs no change.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Chores
    • Improved test infrastructure with finer-grained sharding across unit and integration suites for more targeted coverage.
    • Added additional integration and top-level test groupings to increase test coverage of previously-unsharded areas.
    • Updated test execution logic to better identify and serialize heavy shards, improving CI reliability and efficiency.

@vercel

vercel Bot commented Jun 8, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

2 Skipped Deployments
Project Deployment Actions Updated (UTC)
aidotnet_website Ignored Ignored Preview Jun 8, 2026 12:08pm
aidotnet-playground-api Ignored Ignored Preview Jun 8, 2026 12:08pm

@coderabbitai

coderabbitai Bot commented Jun 8, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Warning

Review limit reached

@ooples, we couldn't start this review because you've reached your PR review rate limit.

More reviews will be available in 47 minutes and 37 seconds. Learn how PR review limits work.

Your organization has run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans include higher PR review limits than trial, open-source, and free plans. In all cases, reviews become available again over time. During sustained high-volume PR review activity, CodeRabbit may temporarily slow when the next review becomes available.

Please see our Fair Usage Limits Policy for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: adf67784-2890-4c4f-b7bd-7fd17ef53d33

📥 Commits

Reviewing files that changed from the base of the PR and between 33f6ab4 and 0f47197.

📒 Files selected for processing (1)
  • .github/workflows/sonarcloud.yml

Walkthrough

Refactors the SonarCloud CI test matrix to split coarse model-family shards into finer letter-range shards, adds new integration and top-level shards, updates related comments, and replaces the $heavyShards list to control serialized xunit collection for the new shard names.

Changes

Test Sharding Refinement

Layer / File(s) Summary
Matrix shard definition expansion
.github/workflows/sonarcloud.yml
NeuralNetworks split into A–L and M–Z; Diffusion expanded to six letter-ranges (A–C, D–I, J–M, N–R, S, T–Z); Generated layers split into A–M and N–Z; added combined Code/Forecast/Segment/Survival shard; added ModelFamily TimeSeries/Activation/Loss; integration tests split into letter-based ranges; new top-level namespaces added (FederatedLearning, Audio/Onnx/Tokenization, Components/Adversarial/EndToEnd).
Comments and cleanup notes
.github/workflows/sonarcloud.yml
Updated disk-space cleanup, pre-test snapshot, and serialized-execution comments to list the new heavy shard names.
Heavy shards serialization list update
.github/workflows/sonarcloud.yml
$heavyShards array replaced with the new fine-grained shard names (Diffusion A–C/D–I/J–M/N–R/S/T–Z, Generated A–M/N–Z, NeuralNetworks A–L/M–Z, Code/Forecast/Segment/Survival) and expanded to include selected integration and top-level shards for serialized xunit collection.

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related issues

Possibly related PRs

  • ooples/AiDotNet#555: Also modifies CI test sharding setup in .github/workflows/sonarcloud.yml.
  • ooples/AiDotNet#1485: Changes Diffusion/ModelFamily shard definitions and heavy-shard serialization logic similar to this PR.

Split by letters, each shard aligned,
CI matrix trimmed and better defined,
Heavy shards listed, collection serialized,
Tests march orderly, no longer surprised. 🚦

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title directly and specifically describes the main changes: splitting OOMing model shards finer and adding coverage for ~780 never-run tests including Document AI, Finance, and integration suite tests.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch ci/shard-split-oom-coverage

Comment @coderabbitai help to get the list of available commands and usage tips.

ooples pushed a commit that referenced this pull request Jun 8, 2026
…M-Z, OOM cause)

PR #1540's first commit added ModelFamily-NeuralNetworks A-L / M-Z shards
to replace the old monolithic Unit-08e catch-all, but forgot to remove
08e itself. 08e was filtering on the same ModelFamilyTests.NeuralNetworks
namespace as A-L+M-Z and only excluded 5 unit-test classes that already
have their own 08a-08d dedicated shards.

Net effect: every paper-scale model in the ModelFamily-NN scaffold was
being instantiated + trained twice per CI run (once under A-L/M-Z, once
under 08e). The second pass started after Adam state was already
allocated and pushed total resident set over the runner's 16 GB
envelope, triggering the OOM that's been blocking PR #1540's CI.

Drop 08e and its $heavyShards / disk-cleanup comment references. A-L
and M-Z (model-class-letter-bucketed and serialized via xunit.runner.json
on the matching $heavyShards rows) provide complete non-redundant
coverage of the same ~85 NN model-family test classes.
@vercel

vercel Bot commented Jun 8, 2026

Copy link
Copy Markdown

Deployment failed with the following error:

Resource is limited - try again in 24 hours (more than 100, code: "api-deployments-free-per-day").

Learn More: https://vercel.com/franklins-projects-02a0b5a0?upgradeToPro=build-rate-limit

@vercel

vercel Bot commented Jun 8, 2026

Copy link
Copy Markdown

Deployment failed with the following error:

Resource is limited - try again in 1 day (more than 100, code: "api-deployments-free-per-day").

Learn More: https://vercel.com/franklins-projects-02a0b5a0?upgradeToPro=build-rate-limit

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
.github/workflows/sonarcloud.yml (1)

678-685: ⚠️ Potential issue | 🟠 Major

Make the PowerShell JSON rewrite resilient: use Add-Member -Force for parallelizeTestCollections / maxParallelThreads

Directly assigning $cfg.parallelizeTestCollections / $cfg.maxParallelThreads after ConvertFrom-Json can throw on a PSCustomObject if those keys are missing in the workflow-produced xunit.runner.json, which would break the intended heavy-shard OOM mitigation. The script already uses Add-Member -Force for preEnumerateTheories, so extend that same safe pattern to these two keys as well.

Proposed fix
              $cfg = Get-Content $runnerJson -Raw | ConvertFrom-Json
-              $cfg.parallelizeTestCollections = $false
-              $cfg.maxParallelThreads = 1
+              $cfg | Add-Member -NotePropertyName parallelizeTestCollections -NotePropertyValue $false -Force
+              $cfg | Add-Member -NotePropertyName maxParallelThreads -NotePropertyValue 1 -Force
               # Don't materialize every [Theory]'s data rows at discovery time —
               # the assembly discovers ~64k cases and pre-enumeration inflates the
               # retained-case baseline before a single test even runs.
               $cfg | Add-Member -NotePropertyName preEnumerateTheories -NotePropertyValue $false -Force
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In @.github/workflows/sonarcloud.yml around lines 678 - 685, The script mutates
the PSCustomObject $cfg by assigning $cfg.parallelizeTestCollections and
$cfg.maxParallelThreads directly which can fail if those properties don't exist;
change those two assignments to use Add-Member -NotePropertyName
parallelizeTestCollections / -NotePropertyName maxParallelThreads with
-NotePropertyValue (false / 1) and -Force (matching how preEnumerateTheories is
added) so the JSON rewrite is resilient when keys are absent, then
ConvertTo-Json and Set-Content as before.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In @.github/workflows/sonarcloud.yml:
- Around line 526-527: Update the troubleshooting comment text that still
references the old shard group names (e.g., "Diffusion S-Z", "Diffusion J-R",
"Generated", "ModelFamily-NN") so it matches the current matrix naming; replace
those old names with the new shard names used in the matrix ("Diffusion S",
"Diffusion T-Z", "Diffusion J-M", "Diffusion N-R", and "Generated Layers
A-M/N-Z") wherever they appear in the sonarcloud workflow comments (the blocks
around the current coverage/dump guidance). Ensure all occurrences in the
workflow comments are synchronized so runbook instructions point to the correct
shard sets.

---

Outside diff comments:
In @.github/workflows/sonarcloud.yml:
- Around line 678-685: The script mutates the PSCustomObject $cfg by assigning
$cfg.parallelizeTestCollections and $cfg.maxParallelThreads directly which can
fail if those properties don't exist; change those two assignments to use
Add-Member -NotePropertyName parallelizeTestCollections / -NotePropertyName
maxParallelThreads with -NotePropertyValue (false / 1) and -Force (matching how
preEnumerateTheories is added) so the JSON rewrite is resilient when keys are
absent, then ConvertTo-Json and Set-Content as before.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 65c81052-a4ad-407c-ad1f-2a9c04ca1f6e

📥 Commits

Reviewing files that changed from the base of the PR and between 1be2ff3 and 33f6ab4.

📒 Files selected for processing (1)
  • .github/workflows/sonarcloud.yml

Comment thread .github/workflows/sonarcloud.yml Outdated
Splits the Diffusion / Generated / NeuralNetworks ModelFamily shards by model
class (A-L / M-Z, Diffusion into 6) so each shard instantiates a bounded subset
instead of all paper-scale models at once — fixes the 16 GB ubuntu-latest OOM
("received a shutdown signal", 0 failures). Also adds matrix coverage for ~780
previously never-run tests (Document AI, Finance, integration suite).

Rebased cleanly onto master; net change is sonarcloud.yml only.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant