Repository navigation
ci(deps): bump actions/github-script from 7 to 8 - #713
Merged
Merged
Conversation
Bumps [actions/github-script](https://github.com/actions/github-script) from 7 to 8. - [Release notes](https://github.com/actions/github-script/releases) - [Commits](actions/github-script@v7...v8) --- updated-dependencies: - dependency-name: actions/github-script dependency-version: '8' dependency-type: direct:production update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com>
Contributor
Author
LabelsThe following labels could not be found: Please fix the above issues or remove invalid values from |
Contributor
|
Important Review skippedBot user detected. To trigger a single review, invoke the You can disable this status message by setting the Comment |
ooples
added a commit
that referenced
this pull request
Jun 30, 2026
…te (#1745) Replaces the interim "gate proactive BF16 on memory pressure" workaround with the real fix: BF16-Adam now keeps the fused fast path instead of dropping to the eager autograd tape, so large models get BOTH the fused speed AND the halved optimizer-state footprint — no tradeoff. Pairs with AiDotNet.Tensors PR #713 (fused bf16 moment kernel + ICompiledTrainingPlan.RequestBf16MomentStorage): - Adam8BitOptimizer implements IFusedOptimizerSpec: in BFloat16 moment-storage mode it maps to the fused Adam kernel with UseBf16Moments=true. The true 8-bit block-quant mode (and adaptive-LR / AMSGrad) still has no fused kernel and correctly falls back to eager. - FusedOptimizerConfig carries UseBf16Moments; TryMapToFusedOptimizerConfig surfaces it; CompiledTapeTrainingStep calls plan.RequestBf16MomentStorage before ConfigureOptimizer so the plan allocates half-size m/v buffers. - ShouldUseBFloat16Optimizer reverts to a plain size threshold — the memory gate existed only to avoid losing the fused path, which no longer happens. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
ooples
pushed a commit
that referenced
this pull request
Jul 1, 2026
- FusedOptimizerConfig: move UseBf16Moments from the primary constructor to an init-only property so Deconstruct arity and positional construction sites are unchanged (only Adam8Bit sets it, now via object initializer); still part of record value equality. - TryStepWithFusedOptimizer: append useBf16Moments after eagerOptimizer instead of inserting it before, so positional call sites aren't shifted (sole caller uses named args). - Directory.Packages.props: document the 0.104.6 -> 0.106.0 bump (Tensors #713 fused bf16 moment kernel) per the file's changelog convention; note 0.106.0 is already published so CI isn't gated on an unreleased dependency. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
ooples
added a commit
that referenced
this pull request
Jul 2, 2026
…ernel, not a gate) (#1745) * perf(training): gate proactive BF16-Adam on real memory pressure (keeps ≥50M models on the fused path) ShouldUseBFloat16Optimizer engaged BF16 moment storage proactively for ANY model with ≥50M parameters. But the BF16/8-bit Adam (Adam8BitOptimizer) is NOT fused-kernel- compatible — TryMapToFusedOptimizerConfig only accepts plain Adam/AdamW/SGD — so selecting it silently drops the ENTIRE model off the compiled fused-training fast path onto the eager autograd tape, ~10x slower per step. Net effect: every ≥50M-param model was training on the slow path even when it fit in memory with room to spare. Measured on ViT-Base (86.5M): the proactive BF16 forced the eager tape at ~5.0 s/step; gating it off (the model fits) keeps it on the fused path at ~3.4 s/step (~1.5x), with no change to small (<50M) models. Fix: only engage proactive BF16 when the fp32 moment state would actually consume a meaningful fraction of available memory (> 25%), mirroring the fits-in-memory guard used for weight streaming. Models that genuinely don't fit still get BF16. The reactive memory ladder (_memoryLeversForced, set on an actual OOM) is untouched and still engages BF16/8-bit on demand. AIDOTNET_BF16_ADAM=1/0 still force/disable explicitly. This is the broad half of the "memory-lever optimizers break fused training" finding in #1743. Refs #1743, #1706. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * perf(training): make BF16-Adam fused-compatible; remove the memory gate (#1745) Replaces the interim "gate proactive BF16 on memory pressure" workaround with the real fix: BF16-Adam now keeps the fused fast path instead of dropping to the eager autograd tape, so large models get BOTH the fused speed AND the halved optimizer-state footprint — no tradeoff. Pairs with AiDotNet.Tensors PR #713 (fused bf16 moment kernel + ICompiledTrainingPlan.RequestBf16MomentStorage): - Adam8BitOptimizer implements IFusedOptimizerSpec: in BFloat16 moment-storage mode it maps to the fused Adam kernel with UseBf16Moments=true. The true 8-bit block-quant mode (and adaptive-LR / AMSGrad) still has no fused kernel and correctly falls back to eager. - FusedOptimizerConfig carries UseBf16Moments; TryMapToFusedOptimizerConfig surfaces it; CompiledTapeTrainingStep calls plan.RequestBf16MomentStorage before ConfigureOptimizer so the plan allocates half-size m/v buffers. - ShouldUseBFloat16Optimizer reverts to a plain size threshold — the memory gate existed only to avoid losing the fused path, which no longer happens. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(training): Adam8Bit BF16 mode maps to fused Adam; block-quant/AMSGrad fall back (#1745) * chore(deps): bump aidotnet.tensors + native packages to 0.106.0 0.106.0 ships the fused bf16 adam moment kernel (tensors #713: adamupdatebf16simd / requestbf16momentstorage) this pr wires into compiledtapetrainingstep, plus the weightregistry dead-owner sweep (tensors #716). * fix(review): bf16-Adam config as init property, param order, changelog - FusedOptimizerConfig: move UseBf16Moments from the primary constructor to an init-only property so Deconstruct arity and positional construction sites are unchanged (only Adam8Bit sets it, now via object initializer); still part of record value equality. - TryStepWithFusedOptimizer: append useBf16Moments after eagerOptimizer instead of inserting it before, so positional call sites aren't shifted (sole caller uses named args). - Directory.Packages.props: document the 0.104.6 -> 0.106.0 bump (Tensors #713 fused bf16 moment kernel) per the file's changelog convention; note 0.106.0 is already published so CI isn't gated on an unreleased dependency. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore(deps): bump aidotnet.tensors + native packages to 0.106.1 --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: franklinic <franklin@ivorycloud.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Rebasing might not happen immediately, so don't worry if this takes some time.
Note: if you make any changes to this PR yourself, they will take precedence over the rebase.
Bumps actions/github-script from 7 to 8.
Release notes
Sourced from actions/github-script's releases.
... (truncated)
Commits
ed59741Merge pull request #653 from actions/sneha-krip/readme-for-v82dc352eBold minimum Actions Runner version in README01e118cUpdate README for Node 24 runtime requirements8b222acApply suggestion from@salmanmkcadc0eeaREADME for updating actions/github-script from v7 to v820fe497Merge pull request #637 from actions/node24e7b7f22update licenses2c81ba0Update Node.js version support to 24.xDependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting
@dependabot rebase.Dependabot commands and options
You can trigger Dependabot actions by commenting on this PR:
@dependabot rebasewill rebase this PR@dependabot recreatewill recreate this PR, overwriting any edits that have been made to it@dependabot mergewill merge this PR after your CI passes on it@dependabot squash and mergewill squash and merge this PR after your CI passes on it@dependabot cancel mergewill cancel a previously requested merge and block automerging@dependabot reopenwill reopen this PR if it is closed@dependabot closewill close this PR and stop Dependabot recreating it. You can achieve the same result by closing it manually@dependabot show <dependency name> ignore conditionswill show all of the ignore conditions of the specified dependency@dependabot ignore this major versionwill close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself)@dependabot ignore this minor versionwill close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself)@dependabot ignore this dependencywill close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself)