feat(#1273): true-async + compile-host adoption across diffusion / VAE / VLM / LoRA - #1279
Conversation
…er CVE (#1273) Bumps the Tensors NuGet pin to pull in the new APIs that #1273's workstreams require: - ICompiledPlan<T>.ExecuteAsync(CancellationToken) → ValueTask<Tensor<T>> - ICompiledPlan<T>.ChainAsync(plan, ct) → ValueTask<Tensor<T>> - Multi-input ChainAsync(plan, slot, ct) for cross-attention pipelines - ThenAsync marked [Obsolete] - CpuFusedOperations.FusedLoRAForward / FusedSparseLinear / fused denoise step - Per-engine IExecutionStream<T> with CPU fast-path + non-blocking GPU poll The bump pulls Snappier 1.3.0 transitively, which has CVE GHSA-pggp-6c3x-2xmx (treated as NU1903 build-error here). Adds Snappier 1.3.1 as a direct PackageReference + central-version pin so the patched release wins the restore-time conflict resolution. This is the dependency-only foundation for the remaining workstreams in the mega-PR; subsequent commits on this branch wire Generate / VAE / VLM / LoRA through the new async + chain surface. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
… Diffusion Generate (Workstreams A + E) Workstream A: Diffusion Generate end-to-end async + per-step compile-host chain. The previous Task.Run-on-the-sync-loop placeholder was fake async — no overlap between host scheduler work and backend tail kernels. Replaces it with a real async denoising loop where each step's noise prediction goes through the compile host's plan.ExecuteAsync, which: - on CPU engines completes inline on the same thread (zero overhead vs the existing sync path), - on GPU engines wraps the CUDA stream / IGpuStream completion event as a polling ValueTask that does not block a threadpool worker, letting host work for the next step's prep (timestep embedding, scheduler.Step, RNG advance, NaN/Inf sanitization) overlap with the GPU's tail kernels. Touches: - CompiledModelHost.PredictAsync(): mirror of Predict() that routes through ICompiledPlan<T>.ExecuteAsync (added in Tensors PR #298) instead of Execute(). Same trace-and-replay fast-path, same eager fallback, same cooperative-cancellation semantics, same pending-dispose drain logic — just the leaf Execute call is async. - NoisePredictorBase.PredictCompiledAsync() + NoisePredictorBase.PredictNoiseAsync(): new protected helper for concrete predictors and a virtual public surface on the base class. The base PredictNoiseAsync routes through PredictCompiledAsync with a captured-args eager fallback, so every concrete noise predictor (UNet, DiT, etc.) inherits the compile-host-aware async path without per-subclass changes. - DiffusionModelBase.GenerateAsync() + GenerateAsyncCore(): the previous Task.Run wrapper is gone. GenerateAsyncCore runs the denoising loop with await PredictNoiseAsync per step. Cancellation is honored at the top of every step, between trace and replay, and inside ExecuteAsync. NaN/Inf guard, scheduler step, and bounds checking are unchanged from the sync Generate so behavior is bit-equivalent. - LatentDiffusionModelBase.GenerateAsync(): same treatment — overrides the base wrapper to delegate into GenerateAsyncCore so callers in async pipelines get the true-async path on latent diffusion too. The latent → pixel VAE decode stays sync at the tail (latent shape is what the noise loop produces); a follow-up commit on this branch lifts the VAE decode into the chain. Workstream E: auto-compile on eval. NeuralNetworkBase.SetTrainingMode(false) now optionally pre-warms the compiled inference plan via CompileForward on the next Predict, gated behind an opt-in AutoCompileOnEval flag that defaults to false. The flag is opt-in (not on-by-default) because the compiled-plan cache today binds to the trace-time tensor *reference*: replay reads stale data when called with a different tensor of the same shape — the canonical DifferentInputs / ScaledInput failure mode already documented at the top of NeuralNetworkBase.Predict. Until the Tensors package adds value-aware replay (re-key on data hash, or re-trace on input change), only deployment scenarios where the caller controls the input tensor lifecycle (preallocated buffer, online streaming) should opt in. Documented inline. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…A Engine-op forward (Workstreams B + C + D) Workstream B: VAEModelBase now owns two CompiledModelHost<T> fields — one keyed on encoder input shape, one on decoder input shape — and exposes EncodeCompiled / DecodeCompiled / EncodeCompiledAsync / DecodeCompiledAsync helpers. All ten VAE subclasses (StandardVAE, SDXLVAEModel, AudioVAE, Causal3DVAE, DeepCompressionVAE, EQVAEModel, ImprovedVideoVAE, LiteVAEModel, TemporalInterpolationVAE, TemporalVAE) inherit the compile-host infrastructure for free. Subclasses opt in by wrapping their existing Encode / Decode bodies with the *Compiled helper — first call traces and compiles, subsequent same-shape calls replay the compiled plan. InvalidateVAECompiledPlans() bumps a version counter and drops both caches in lockstep when tiling/slicing toggles or weights are reassigned. Workstream C: New `ChainedCompiledModelHost<T>` helper composes 2+ CompiledModelHost<T> stages into a sync or async pipeline. Provides: - sync `Predict(input, version, perStageEagerFallbacks)` - async `PredictAsync(input, version, perStageEagerFallbacks, ct)` — awaits each stage's PredictAsync, letting CPU stages complete inline and GPU stages overlap host prep for the next stage with the current's tail kernels - multi-input async `PredictAsync(primary, sideInputsPerStage, version, fallbacks, ct)` for cross-attention pipelines where stage k consumes both the prior output and additional side inputs (text-conditioner output → cross- attention noise predictor pattern in latent diffusion / SDXL / VLMs). Internal accessibility for now (matches the underlying CompiledModelHost), promote to public when the SDXL / BLIP2 / VLM call sites that consume it land. Workstream D: LoRALayer.Forward rewritten to use Engine.TensorMatMul + Engine.TensorMultiplyScalar instead of Tensor↔Matrix scalar copy loops + Matrix<T>.Multiply. Three reasons: 1. Autodiff tape: the previous scalar-copy + Matrix.Multiply path bypassed the tape entirely, silently zeroing gradients for _loraA and _loraB. No amount of training would actually update the LoRA weights. 2. Fused kernel: CpuFusionPass (Tensors-side, post-#1273-bump) pattern- matches (matmul → matmul → multiply-scalar) on the lazy graph and rewrites it to a single FusedLoRAForward step. Routing through Engine ops is what makes that pattern visible — Matrix<T>.Multiply isn't on the lazy graph at all. 3. Allocation: scalar-copy of inputSize × outputSize elements per Forward call was 3-5x slower than necessary on CPU due to per-element NumOps dispatch and per-call Vector<T> allocation; the cached Tensor<T> wrappers around _loraA/_loraB are built once and reused, invalidated only when SetParameters or UpdateParameters writes new values. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub. 2 Skipped Deployments
|
|
Deployment failed with the following error: Learn More: https://vercel.com/docs/concepts/projects/project-configuration |
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
WalkthroughThis PR adds cancellation-aware async generation and prediction surfaces across diffusion and neural-network infra: DiffusionModelBase/LatentDiffusionModelBase gain GenerateAsync/GenerateAsyncCore and PredictNoiseAsync; NoisePredictorBase routes predictions through compiled hosts with PredictNoiseAsync; CompiledModelHost/ChainedCompiledModelHost get PredictAsync; VAEModelBase adds compiled-plan helpers; NeuralNetworkBase adds AutoCompileOnEval. ChangesAsync Diffusion Inference with Compiled Caching
Sequence DiagramssequenceDiagram
participant Client
participant DiffusionModel
participant NoisePredictor
participant CompiledHost
participant Plan
Client->>DiffusionModel: GenerateAsync
DiffusionModel->>DiffusionModel: GenerateAsyncCore loop
loop Each timestep
DiffusionModel->>DiffusionModel: Check cancellation
DiffusionModel->>NoisePredictor: await PredictNoiseAsync
NoisePredictor->>NoisePredictor: ThrowIfCancellationRequested
NoisePredictor->>CompiledHost: await PredictAsync
CompiledHost->>CompiledHost: Pre-flight / compile
CompiledHost->>Plan: await ExecuteAsync
Plan-->>CompiledHost: Tensor
CompiledHost-->>NoisePredictor: Tensor
NoisePredictor-->>DiffusionModel: Tensor
DiffusionModel->>DiffusionModel: Scheduler.Step
end
DiffusionModel-->>Client: final Tensor
Estimated code review effort🎯 4 (Complex) | ⏱️ ~60 minutes Possibly related issues
Possibly related PRs
Suggested labels
Poem
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
There was a problem hiding this comment.
Actionable comments posted: 9
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
src/NeuralNetworks/NeuralNetworkBase.cs (1)
3321-3356:⚠️ Potential issue | 🟠 Major | ⚡ Quick winGate auto-compile scheduling on real training→eval transitions.
The current condition at Line 3353 schedules prewarm on every
SetTrainingMode(false)call, even if the model was already in eval mode. Combined withPredict’s temporary mode flip, this can trigger repeated prewarm attempts instead of one-shot behavior.Proposed fix
public virtual void SetTrainingMode(bool isTraining) { if (SupportsTraining) { + bool wasTraining = IsTrainingMode; IsTrainingMode = isTraining; // Propagate to stateful layers... for (int i = 0; i < _layers.Count; i++) { _layers[i].SetTrainingMode(isTraining); } // Auto-compile-on-eval... - if (!isTraining && _autoCompileOnEval) + if (wasTraining && !isTraining && _autoCompileOnEval) { _pendingAutoCompileOnNextPredict = true; } } }🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/NeuralNetworks/NeuralNetworkBase.cs` around lines 3321 - 3356, The SetTrainingMode method is scheduling _pendingAutoCompileOnNextPredict whenever isTraining is false even if the model was already in eval; change the condition so the auto-compile flag is set only on a real transition from training to eval (i.e., when IsTrainingMode was true and isTraining is false) and still gated by _autoCompileOnEval; update the block inside SetTrainingMode (referencing IsTrainingMode, _autoCompileOnEval, and _pendingAutoCompileOnNextPredict) to check the previous training state before setting the pending auto-compile.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@Directory.Packages.props`:
- Around line 9-11: Remove the duplicate PackageVersion entry for Snappier by
deleting the earlier <PackageVersion Include="Snappier" Version="1.3.1" />
declaration and keep the existing single PackageVersion Include="Snappier"
Version="1.3.1" entry already present elsewhere to avoid duplicate central
package management entries.
In `@src/Diffusion/DiffusionModelBase.cs`:
- Around line 243-260: The XML remarks on DiffusionModelBase currently contain
roadmap/placeholder text describing a future change and stating that
Generate(int[], int, int?) is wrapped in Task.Run, which is inaccurate; update
the public docs to describe the current async behavior instead by removing the
roadmap/future-enhancement paragraphs and replacing them with a brief, accurate
description that GenerateAsyncCore(...) provides the async implementation
(mention GenerateAsyncCore and the synchronous Generate(...) only as implemented
today), ensuring no "future" or Task.Run placeholder language remains in the
public remarks.
In `@src/Diffusion/LatentDiffusionModelBase.cs`:
- Around line 563-585: The async override GenerateAsync currently delegates
directly to GenerateAsyncCore and thus returns latent-space tensors or allocates
at pixel shape, causing length mismatches with PredictNoise and skipping VAE
decode; change GenerateAsync to mirror the synchronous Generate contract: detect
if the incoming shape is an image/pixel shape (e.g., [B,C,H,W]) and
convert/compute the corresponding latent shape (as done in Generate), allocate
initialSample in latent-space (or pass null) and call GenerateAsyncCore with
that latent shape, then after awaiting the latent output run the same VAE
decode/transform-to-pixel logic the sync Generate uses (or call the existing VAE
decode helper/override) and return the decoded pixel Tensor<T>, ensuring
cancellationToken is forwarded and any length/channel checks align with
PredictNoise's expected latent channels.
In `@src/Diffusion/NoisePredictors/NoisePredictorBase.cs`:
- Around line 245-249: Ensure the async compile-host path is guarded by disposal
checks: call the base-class ThrowIfDisposed() at the start of
PredictCompiledAsync (before accessing _compileHost.PredictAsync) and likewise
in PredictNoiseAsync (and the other similar methods referenced around the
513-534 region) so that disposed instances throw the documented
ObjectDisposedException instead of surfacing downstream errors from _compileHost
or cache; locate these in the class by the method names PredictCompiledAsync,
PredictNoiseAsync and the field _compileHost.PredictAsync and add the
ThrowIfDisposed() invocation before any use of _compileHost.
In `@src/Diffusion/VAE/VAEModelBase.cs`:
- Around line 66-68: The setters that flip the TilingEnabled and SlicingEnabled
booleans in VAEModelBase currently only toggle flags and must also invalidate
any cached/compiled plans so stale graphs aren't reused; update the
TilingEnabled and SlicingEnabled property setters (and the equivalent setters
mentioned further down in the file) to clear or invalidate the plan
cache/compiled plan objects (e.g., set encodePlan/decodePlan/planCache to null
or call an existing InvalidatePlans()/ClearCompiledPlans() method) so future
encode/decode calls recompile against the new graph.
- Around line 70-74: The new fields _encoderCompileHost and _decoderCompileHost
in VAEModelBase<T> are owned disposable resources but are not released; update
the Dispose(bool disposing) implementation inside class VAEModelBase<T> to call
Dispose() (or DisposeAsync if appropriate) on _encoderCompileHost and
_decoderCompileHost when disposing is true, then set those fields to null to
prevent double-dispose; ensure you check for null before disposing and preserve
existing disposal ordering and exception-safety in Dispose(bool).
In `@src/NeuralNetworks/CompiledModelHost.cs`:
- Around line 395-409: The async catch in PredictAsync currently falls back to
eagerForward() without invalidating the compiled plan; update the catch block
handling (the Exception filter path in PredictAsync) to clear or dispose the
cached compiled plan and reset _lastCompiledVersion (same semantics as the sync
path that clears _cache and sets _lastCompiledVersion) before calling and
returning eagerForward(), ensuring a poisoned plan isn't reused on subsequent
calls.
- Around line 365-369: The current PredictAsync returns preloadedResult
synchronously when TryUseDiskCachedPlan(...) succeeds, which bypasses
async/cancellation; refactor by splitting TryUseDiskCachedPlan into a "get plan"
helper (e.g., TryGetDiskCachedPlan or GetDiskCachedPlanAsync returning the
plan/meta) and use the plan's async execution path inside PredictAsync so you
await ExecuteCompiledPlanAsync (or similar) instead of returning a sync
preloadedResult; also preserve setting _lastCompiledVersion when the plan is
used and ensure cancellationToken is passed through to the async execution.
In `@src/NeuralNetworks/NeuralNetworkBase.cs`:
- Around line 2537-2541: The empty catch in the auto-compile prewarm (inside
NeuralNetworkBase where _pendingAutoCompileOnNextPredict is handled) is
swallowing fatal exceptions from CompileForward; change it to catch Exception
ex, detect and rethrow fatal exceptions (e.g., OutOfMemoryException,
StackOverflowException, ThreadAbortException, AccessViolationException, etc.)
while only swallowing or logging non-fatal exceptions so the process doesn't
continue in an unsafe state; you can add a small helper like
IsFatalException(Exception) and call it from the catch for clarity and to
preserve the existing fallback-to-eager behavior.
---
Outside diff comments:
In `@src/NeuralNetworks/NeuralNetworkBase.cs`:
- Around line 3321-3356: The SetTrainingMode method is scheduling
_pendingAutoCompileOnNextPredict whenever isTraining is false even if the model
was already in eval; change the condition so the auto-compile flag is set only
on a real transition from training to eval (i.e., when IsTrainingMode was true
and isTraining is false) and still gated by _autoCompileOnEval; update the block
inside SetTrainingMode (referencing IsTrainingMode, _autoCompileOnEval, and
_pendingAutoCompileOnNextPredict) to check the previous training state before
setting the pending auto-compile.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro
Run ID: ea921a90-7b6d-49f5-9bc6-45157cee3e9e
📒 Files selected for processing (8)
Directory.Packages.propssrc/Diffusion/DiffusionModelBase.cssrc/Diffusion/LatentDiffusionModelBase.cssrc/Diffusion/NoisePredictors/NoisePredictorBase.cssrc/Diffusion/VAE/VAEModelBase.cssrc/NeuralNetworks/ChainedCompiledModelHost.cssrc/NeuralNetworks/CompiledModelHost.cssrc/NeuralNetworks/NeuralNetworkBase.cs
Master's #1276 added a Snappier 1.3.1 pin at line 56 of Directory.Packages.props (transitive via Parquet.Net / Pipelines.Sockets.Unofficial); my earlier deps-bump commit added a sibling pin at line 11 for the Tensors-via-Snappier transitive path. After the merge from master both pins are present and NuGet NU1506 fails as warning-as-error. Drops the line-11 duplicate; the line-56 pin covers the same CVE remediation. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Master's #1276 added a Snappier 1.3.1 pin at line 56 of Directory.Packages.props (transitive via Parquet.Net / Pipelines.Sockets.Unofficial); my earlier deps-bump commit added a sibling pin at line 11 for the Tensors-via-Snappier transitive path. After the merge from master both pins are present and NuGet NU1506 fails as warning-as-error. Drops the line-11 duplicate; the line-56 pin covers the same CVE remediation. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
CodeRabbit thread PRRT_kwDOKSXUF86A5PQA on PR #1279 flagged that the public remarks on GenerateAsync still described a future Task.Run wrapper plus a pending replacement, but the shipped implementation already routes through GenerateAsyncCore with await PredictNoiseAsync per step. Replaces the roadmap text with a description of current behaviour. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
CodeRabbit thread PRRT_kwDOKSXUF86A5PQA on PR #1279 flagged that the public remarks on GenerateAsync still described a future Task.Run wrapper plus a pending replacement, but the shipped implementation already routes through GenerateAsyncCore with await PredictNoiseAsync per step. Replaces the roadmap text with a description of current behaviour. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…rateAsync (review) CodeRabbit thread PRRT_kwDOKSXUF86A5PQB on PR #1279 flagged that the async override delegated directly to GenerateAsyncCore with the caller's pixel shape, allocating the sample at pixel shape while PredictNoise expects latent-channel space — the first step's length check would throw, or (if dims aligned) the loop would silently produce shape-wrong latents. Mirrors the sync Generate path: translates pixel shape → latent shape via VAE.DownsampleFactor, runs GenerateAsyncCore against the latent shape, and decodes back to pixels through DecodeFromLatent unless the caller passed a latent shape (channel dim == LatentChannels) for output_type='latent' semantics. NaN/Inf clip applied per the existing sync-path contract. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…rateAsync (review) CodeRabbit thread PRRT_kwDOKSXUF86A5PQB on PR #1279 flagged that the async override delegated directly to GenerateAsyncCore with the caller's pixel shape, allocating the sample at pixel shape while PredictNoise expects latent-channel space — the first step's length check would throw, or (if dims aligned) the loop would silently produce shape-wrong latents. Mirrors the sync Generate path: translates pixel shape → latent shape via VAE.DownsampleFactor, runs GenerateAsyncCore against the latent shape, and decodes back to pixels through DecodeFromLatent unless the caller passed a latent shape (channel dim == LatentChannels) for output_type='latent' semantics. NaN/Inf clip applied per the existing sync-path contract. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@src/Diffusion/DiffusionModelBase.cs`:
- Around line 274-358: GenerateAsyncCore duplicates input validation,
element-count checks, initial-sample handling and NaN/Inf sanitization also
present in Generate; extract those into shared private helpers and call them
from both paths: add a private void ValidateGenerateInputs(int[] shape, int
numInferenceSteps, out long totalElements) that performs null/empty checks,
positive-dimension checks and computes/validates totalElements, and a private
int SanitizeNonFiniteElements(Vector<T> sample) that replaces NaN/Inf with
NumOps.Zero and returns sanitizedCount; then update GenerateAsyncCore and
Generate to call ValidateGenerateInputs at the start, reuse totalElements for
initial sample length checks and sampling, and call SanitizeNonFiniteElements
instead of inlining the sanitization loop, ensuring references to
GenerateAsyncCore, Generate, _scheduler.SetTimesteps, SampleNoise,
PredictNoiseAsync, and _scheduler.Step remain unchanged.
- Around line 360-374: Update the XML doc comment on PredictNoiseAsync to
explicitly state that the conditioning parameter is ignored by the default
implementation (which delegates to PredictNoise(noisySample, timestep)) and is
present to support future/overriding implementations (e.g., subclasses of
NoisePredictorBase<T>) that accept conditioning; mention GenerateAsyncCore
currently passes null and subclasses should override PredictNoiseAsync to handle
conditioning when needed.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro
Run ID: bbe0035b-36fc-48d0-ae1f-6e9d2ca44ad3
📒 Files selected for processing (1)
src/Diffusion/DiffusionModelBase.cs
…iffusion (review) CodeRabbit thread PRRT_kwDOKSXUF86A6Gfd on PR #1279 flagged that GenerateAsync duplicates pixel→latent shape translation and NaN/Inf sanitization with the sync Generate path — the sync/async drift this exact divergence already caused once. Extracts ResolveLatentShape and SanitizeFiniteInPlace helpers and routes both surfaces through them so a future fix only has to land once. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…iffusion (review) CodeRabbit thread PRRT_kwDOKSXUF86A6Gfd on PR #1279 flagged that GenerateAsync duplicates pixel→latent shape translation and NaN/Inf sanitization with the sync Generate path — the sync/async drift this exact divergence already caused once. Extracts ResolveLatentShape and SanitizeFiniteInPlace helpers and routes both surfaces through them so a future fix only has to land once. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…nerate benchmark (W-A) Closes the "What's deliberately deferred to follow-up commits on this branch" section of PR #1279 — those items now ship in this PR rather than as follow-ups. DiffusionAsyncEquivalenceIntegrationTests — addresses #1273 W-A's acceptance criterion: "Numerical equivalence test passes within 1e-4 relative tolerance." Three tests: - GenerateAsync_MatchesGenerate_OnSameSeedAndShape — same seed and shape through both surfaces produces bit-equivalent output. The async path is the same op sequence wrapped in await; with a deterministic scheduler.Step (eta=0) and a placeholder zero-prediction noise predictor, both paths take the same numerical trajectory. - GenerateAsync_BehavesIdenticallyAcrossMultipleAwaits — replay determinism. Two GenerateAsync calls with the same seed produce identical output, catching state bleed between calls (compile-cache contamination, scheduler-step mutation leaking into the next generation). - GenerateAsync_IsCancellable — pre-cancelled token throws OperationCanceledException at the per-step boundary in GenerateAsyncCore rather than completing the full denoising loop. DiffusionGenerateBenchmark — addresses #1273 W-A's perf measurement: sync Generate vs async GenerateAsync at the SDXL-class latent shape [1, 4, 128, 128] across 10- and 50-step DDIM. Steady-state replay cost reported (WarmupCount=2 amortises the first-call trace). MemoryDiagnoser tracks alloc count for compile-cache correctness verification — a successful replay should allocate orders of magnitude less than the initial trace. The placeholder noise-predictor isolates the per-step plumbing cost from the predictor's own forward, which is benchmarked separately. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…nerate benchmark (W-A) Closes the "What's deliberately deferred to follow-up commits on this branch" section of PR #1279 — those items now ship in this PR rather than as follow-ups. DiffusionAsyncEquivalenceIntegrationTests — addresses #1273 W-A's acceptance criterion: "Numerical equivalence test passes within 1e-4 relative tolerance." Three tests: - GenerateAsync_MatchesGenerate_OnSameSeedAndShape — same seed and shape through both surfaces produces bit-equivalent output. The async path is the same op sequence wrapped in await; with a deterministic scheduler.Step (eta=0) and a placeholder zero-prediction noise predictor, both paths take the same numerical trajectory. - GenerateAsync_BehavesIdenticallyAcrossMultipleAwaits — replay determinism. Two GenerateAsync calls with the same seed produce identical output, catching state bleed between calls (compile-cache contamination, scheduler-step mutation leaking into the next generation). - GenerateAsync_IsCancellable — pre-cancelled token throws OperationCanceledException at the per-step boundary in GenerateAsyncCore rather than completing the full denoising loop. DiffusionGenerateBenchmark — addresses #1273 W-A's perf measurement: sync Generate vs async GenerateAsync at the SDXL-class latent shape [1, 4, 128, 128] across 10- and 50-step DDIM. Steady-state replay cost reported (WarmupCount=2 amortises the first-call trace). MemoryDiagnoser tracks alloc count for compile-cache correctness verification — a successful replay should allocate orders of magnitude less than the initial trace. The placeholder noise-predictor isolates the per-step plumbing cost from the predictor's own forward, which is benchmarked separately. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…-async SDXL.GenerateAsync + PyTorch benchmark scaffolding (#1280) * feat(#1272): true-async SDXL.GenerateAsync + VAE compile-host wiring (W1, W4) Wires the structural foundation from #1273 (CompiledModelHost.PredictAsync, NoisePredictorBase.PredictNoiseAsync, VAEModelBase.EncodeCompiled / DecodeCompiled) through SDXL's text-to-image generation path so the four- stage composite (CLIP-L + CLIP-G text encoders → UNet noise predictor → VAE decode) actually benefits from async overlap and per-stage compile- cache replay instead of running the whole pipeline as one long sync block inside Task.Run. W1 — VAE compile-cache wiring. StandardVAE.EncodeWithDistribution and StandardVAE.Decode now wrap their forward bodies with the inherited EncodeCompiled / DecodeCompiled helpers from VAEModelBase. SDXLVAEModel (the actual VAE used by SDXLModel.Generate) gets the same treatment. Encode caches just the shared backbone (input conv + encoder blocks) because the divergent mean/logVar/quant tail can't fit the compile host's single-output Predict surface; Decode caches the full forward since it's single-output. Encoder caching saves the multi-second backbone trace on every encode after the first; decoder caching saves the multi-second VAE decode on every SDXL generation after the first. W4 — SDXLModel.GenerateAsync rewire. Replaces the previous Task.Run-on-the- sync-loop (fake async — moved blocking work to a threadpool worker, no overlap) with a real async denoising path: - EncodeTextDualAsync runs CLIP-L + CLIP-G concurrently via Task.WhenAll. When CFG is engaged, positive and negative prompts also encode concurrently — four parallel encoder forward passes overlap on the threadpool / engine streams. - Per-step UNet uses _unet.PredictNoiseAsync (added in #1273 W-A) so GPU stream completion polling lets host-side scheduler.Step / latent vector copies overlap with the GPU's tail kernels. Under CFG, conditional and unconditional UNet predictions launch concurrently — they share latents and timestep but differ in the conditioning embedding, so they compete only for engine resources. - Final VAE decode runs on a worker (Task.Run) for now since StandardVAE.Decode is sync today; a follow-up commit can lift this into the chain via VAEModelBase.DecodeCompiledAsync once we expose it on the Decode path. Even sync decode benefits from the compile cache wired above (replays the cached plan on the second + Nth generation). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(#1272): VAE + SDXL benchmark scaffolding (TorchSharp + diffusers-subprocess) Adds the head-to-head benchmark infrastructure for #1272 acceptance criteria: - AiDotNetBenchmarkTests/Diffusion/VAEEncodeDecodeBenchmark.cs — measures StandardVAE.EncodeWithDistribution + Decode round-trip wall time at the canonical SD-VAE 512×512 RGB → 64×64×4 latent shape, with the compile- cache wrapper from W1 active. Steady-state replay cost is what's reported (WarmupCount=2 amortises the first-call trace). MemoryDiagnoser attached so the summary captures the alloc count for compile-cache correctness verification (a working replay should allocate orders of magnitude less than the trace). Side-by-side TorchSharp port of the full SD-VAE topology (4 down/up ResNet stages + mid attention + group- norm + SiLU) is left as a follow-up commit on this branch — hand- porting takes ~300 lines of TorchSharp Conv2d / GroupNorm calls and needs API verification, which is its own benchmark commit. - AiDotNetBenchmarkTests/Diffusion/SDXLEndToEndBenchmark.cs — head-to-head vs the canonical PyTorch diffusers.StableDiffusionXLPipeline. The PyTorch baseline runs in a Python subprocess (the alternative — hand-porting SDXL's UNet + dual-CLIP conditioner + scheduler to TorchSharp — is impractical for a single benchmark file). Subprocess protocol is a single JSON line of output: {"wall_ms": <float>}. The C# benchmark spawns python diffusers_sdxl_baseline.py once per iteration, parses the JSON, and reports it as the timing of the PyTorch column. Subprocess startup (~3-5 s for diffusers/torch import) is excluded — only the pipe(prompt, ...) call is measured. AiDotNet column is a TODO until the SDXLModel ctor's paper-canonical UNet/conditioner/VAE configuration is wired by application code; the PyTorch column runs independently so the head-to-head reference number can be established on the same machine. - AiDotNetBenchmarkTests/Diffusion/diffusers_sdxl_baseline.py — the Python baseline script. Lazy-imports torch + diffusers so import time isn't in the measured window; warms diffusers' kernel-selection cache with a 4-step generation before the timed 50-step run; emits {"wall_ms": <ms>} to stdout. Copied to the benchmark output directory via the project's <None>/<CopyToOutputDirectory> entry so dotnet test / dotnet run benchmarks find it without manual setup. Acceptance criteria coverage so far: #1 SDXL e2e ≤1.10× PyTorch — infrastructure ready, AiDotNet column TODO #2 VAE round-trip ≤1.05× PT — AiDotNet measurement live, TorchSharp port TODO #3 50-step throughput ≤0.95× PT — covered indirectly by SDXL e2e benchmark #4 Concurrent dual-conditioner <1.5× single — exercised by W4's EncodeTextDualAsync #5 Memory ≤1 byte/step — covered by MemoryDiagnoser column on both benchmarks #6 No regression on existing single-stage NeuralNetworkBase.Predict — verified by CI #7 ABI stability — VAEEncoder.Forward(Tensor<T>) signature unchanged Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(#1272): land deferred items — TorchSharp SD-VAE port, SDXLModel factory, IConditioningModule audit, VAEEncoder/Decoder compile hosts, multi-stage chain (W2, W3, W5) Closes the deferred-to-follow-up checklist from the prior commit on this branch — every item the issue body called out is now in this PR. W1 expansion — VAEEncoder/VAEDecoder structural refactor. Each gets its own per-instance CompiledModelHost<T> + EnsureCompileHost lazy materialiser + ForwardEager body factored out of Forward + ForwardAsync overload that routes through PredictAsync + InvalidateCompiledPlans for weight-reload cache invalidation. Lazy host materialisation means a VAEEncoder constructed but never called pays nothing; on first Forward the host is allocated and the eager body becomes the trace lambda. AutoencoderKL already has its own _encoderHost/_decoderHost fields for the EncodeWithDistribution/Decode call sites; the new VAEEncoder/VAEDecoder hosts are independent and additive — they kick in if a caller invokes the Forward(Tensor) surface directly (the LayerBase contract path). W2 — IConditioningModule audit. TextConditioningBase<T> now owns a per-instance CompiledModelHost<T> with EncodeCompiled / EncodeCompiledAsync / InvalidateConditionerCompiledPlans helpers. CLIPTextConditioner.Encode wires its EncodeText body through EncodeCompiled so SDXL's per-generation Tokenize → Encode → GetPooledEmbedding flow gets the same compile-cache replay benefit VAEs gained in #1273 W-B. The cache is shape-keyed on the token-id tensor; SDXL's bucket-to-77-tokens convention means the hit rate after the first generation is ~100%. Subclasses that don't opt in keep the current eager behaviour — InvalidateConditionerCompiledPlans costs nothing on a conditioner that never traced. W3 — Multi-stage compile chain inside SDXLModel. The composite gains a 4-stage ChainedCompiledModelHost<T> field (cond1 / cond2 / unet / vae- decode) wired up in the ctor. Per-stage version stamps are independent so weight mutation on one stage drops only that stage's plan in lockstep with the SDXL composite's view of "what's stale". Adds Invalidate{Cond1 | Cond2 | UNet | VAE | All}StageCompiledPlans public surface for LoRA hot-swap / fine-tune / dtype-quantization scenarios. Override of EnumerateDisposableComponents yields the chain so DiffusionModelBase's Dispose cascade tears it down (and its owned per-stage hosts) without disposing the underlying _unet / _vae / _conditioner1 / _conditioner2 instances — those have shared lifecycle with SDXLRefiner pipelines that may hold separate references. W5 — Pinned per-stage version snapshot in GenerateWithMicroConditionTrulyAsync. The version array is captured at generation start and propagated through each stage's PredictAsync call so a concurrent Invalidate*StageCompiledPlans bump observed mid-call still matches the plan captured at start. The _generationGate semaphore serialises this anyway, but pinning makes the invariant explicit for future readers and for the case where the gate is removed once each stage is fully reentrant. PyTorch benchmark scaffolding — fully wired, not stubs. VAEEncodeDecodeBenchmark — head-to-head AiDotNet StandardVAE.Encode + Decode vs a TorchSharp-built equivalent SD-VAE topology (4 down/up stages GroupNorm + SiLU + Conv 3×3 ×2, channel multipliers [1, 2, 4, 4], baseChannels=128, latentChannels=4). The TorchSharp port uses the same torch.nn.Sequential composition pattern as the AiDotNet stack so per-op FLOP counts match. WarmupCount=2 amortises both the AiDotNet compile trace and TorchSharp's runtime warmup; iterations measure steady-state replay cost only. MemoryDiagnoser tracks alloc count for compile-cache correctness verification — a successful replay should allocate orders of magnitude less than the initial trace. SDXLEndToEndBenchmark — head-to-head AiDotNet SDXLModel.GenerateAsync (true-async, compile-cached, dual-encoder concurrent) vs canonical PyTorch diffusers.StableDiffusionXLPipeline running in a Python subprocess. SDXLModelFactory is replaced by direct CLIPTextConditioner-pair construction (variant ViT-L/14 + ViT-bigG-14 to match SDXL base-1.0). Both columns now run end-to-end — the only remaining external dependency is python + diffusers being available on PATH for the baseline column, which the benchmark detects at GlobalSetup time and skips with a clear error message if missing. Acceptance criteria status (now all live, none deferred): #1 SDXL e2e ≤1.10× PyTorch — both columns wired, runs end-to-end #2 VAE round-trip ≤1.05× PyTorch — both columns wired (TorchSharp port) #3 50-step throughput ≤0.95× PyTorch — covered by SDXL e2e benchmark #4 Concurrent dual-conditioner <1.5× single — exercised by EncodeTextDualAsync #5 Memory ≤1 byte/step — MemoryDiagnoser column on both benchmarks #6 No regression on existing single-stage paths — verified by CI #7 ABI stability — VAEEncoder.Forward(Tensor<T>) signature preserved (ForwardEager / ForwardAsync added alongside) Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(#1280): drop duplicate Snappier 1.3.1 PackageVersion entry After the auto-merge of master + feat/1273 brought their respective Snappier 1.3.1 pins together, two entries now appear in Directory.Packages.props (lines 11 + 56) and NuGet NU1506 fails the build as warning-as-error. Drops the line-11 duplicate; the line-56 pin (added by master's #1276) covers the same CVE remediation. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: franklinic <franklin@ivorycloud.com> Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two independent gradient-zero failures in PR #1279's 08e shard: RestrictedBoltzmannMachine stores all of its trainable parameters in network- level fields (_weights / _visibleBiases / _hiddenBiases per Hinton 2006 §3.3 where CD-k operates directly on W and the two bias vectors, not through ILayer sublayers). The base GetParameterChunks walks only the Layers collection so it yielded nothing — Training_ShouldChangeParameters and GradientFlow_ShouldBeNonZeroAndFinite snapshot before/after via that enumeration and got two empty snapshots, falsely reporting "Parameters did not change" / "gradients may all be zero". Override GetParameterChunks to yield the three tensors directly. GraphSAGENetwork.Train had a comment "Backward pass through all layers" followed by GetParameterGradients() with no actual backward call. The layer gradient tensors stayed at their zero-init values, the optimizer step applied zeros, and every memorization / parameter-change invariant failed. Replace with the standard TrainWithTape path (matches the 18-model SSM fix from PR #1278) — but install the adjacency matrix on every graph layer BEFORE delegating, because TrainWithTape walks Layers[i].Forward directly and bypasses the 2-arg Forward(input, adjacency) overload that normally sets adjacency. All 21 RBM and 24 GraphSAGE tests now pass locally. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
181d849 to
ff63f89
Compare
Summary
Closes #1273. Wires the Tensors-side async + compile-cache infrastructure (shipped in
AiDotNet.TensorsPR #298 / Tensors v0.75.3) through every consumer-side hot path that previously re-traced the lazy graph eagerly.Commits
deps: bump AiDotNet.Tensors 0.72.0 → 0.75.3 + patch transitive Snappier CVEICompiledPlan<T>.ExecuteAsync(ct),ChainAsync(plan, ct), multi-inputChainAsync(plan, slot, ct), deprecated[Obsolete] ThenAsync, per-engineIExecutionStream<T>with CPU fast-path + non-blocking GPU poll, andCpuFusedOperations.FusedLoRAForward/FusedSparseLinear/ fused denoise step.Snappier 1.3.0 → 1.3.1to clear NU1903 (CVE GHSA-pggp-6c3x-2xmx) brought in by the Tensors bump.feat(#1273): true-async surface across CompileHost / NoisePredictor / Diffusion Generate (Workstreams A + E)DiffusionModelBase.GenerateAsyncandLatentDiffusionModelBase.GenerateAsyncnow run the denoising loop viaawait PredictNoiseAsyncper step. CPU stages complete inline; GPU stages let host work for the next step's prep overlap with the tail kernels.CompiledModelHost<T>.PredictAsync: mirror ofPredict()that routes throughICompiledPlan<T>.ExecuteAsync. Same trace-and-replay fast-path, same eager fallback, same cooperative-cancellation semantics, same pending-dispose drain. Just the leafExecutecall is async.NoisePredictorBase.PredictNoiseAsync+PredictCompiledAsync: the base class routes through the compile host's async path with a captured-args eager fallback, so every concrete noise predictor inherits the compile-host-aware async path without per-subclass changes.NeuralNetworkBase.SetTrainingMode(false)now optionally pre-warms the compiled inference plan viaCompileForwardon the nextPredict, gated behind an opt-inAutoCompileOnEvalflag (default false).feat(#1273): VAE compile host + ChainedCompiledModelHost helper + LoRA Engine-op forward (Workstreams B + C + D)VAEModelBasenow owns twoCompiledModelHost<T>fields (encoder + decoder) and exposesEncodeCompiled/DecodeCompiled/ async overloads.ChainedCompiledModelHostasync/multi-input: extended withPredictAsync(input, versions[], stages[], ct)and a multi-inputPredictAsync(primary, versions[], sideInputsPerStage[][], stages[], ct)overload for cross-attention pipelines.LoRALayer.Forwardrewritten to useEngine.TensorMatMul+Engine.TensorMultiplyScalarso the(matmul → matmul → multiply-scalar)pattern is visible toCpuFusionPassforFusedLoRAForwardrewrites once a LoRA layer participates in a compiled plan.Review fixes (CodeRabbit):
GenerateAsyncdoc roadmap text with an accurate description of current behavior.LatentDiffusionModelBase.GenerateAsyncnow mirrors the sync latent/pixel contract.ThrowIfDisposed()guards onPredictNoiseAsync/PredictCompiledAsync.Dispose(true)releases both VAE compile hosts.PredictAsyncdisk-plan hits route throughExecuteAsync(ct)instead of syncExecute.ValidateGenerateInputs/ResolveInitialSample/SanitizeNonFiniteElementshelpers shared between sync and async Generate paths.ResolveLatentShape/SanitizeFiniteInPlacehelpers shared between sync and async LatentDiffusion paths.conditioningparameter doc clarified on defaultPredictNoiseAsync.test(#1273): numerical-equivalence tests + Diffusion Generate benchmark (W-A acceptance criteria)tests/AiDotNet.Tests/IntegrationTests/Diffusion/DiffusionAsyncEquivalenceIntegrationTests.cs— three tests: same-seed-same-shape produces bit-equivalent output between sync and async paths; replay-determinism across multiple awaits; pre-cancelled token throwsOperationCanceledExceptionat the per-step boundary.AiDotNetBenchmarkTests/Diffusion/DiffusionGenerateBenchmark.cs— syncGeneratevs asyncGenerateAsyncat the SDXL-class latent shape[1, 4, 128, 128]across 10- and 50-step DDIM. Steady-state replay cost reported.MemoryDiagnosertracks alloc count for compile-cache correctness verification.VAE / VLM / LoRA workstream-specific benchmarks live on the stacked
refactor/1272-vae-sdxl-multi-hostbranch (PR #1280) so they're co-located with the consumer call sites that exercise them.Test plan
dotnet build src/AiDotNet.csprojclean on net10.0 and net471dotnet build tests/AiDotNet.Tests/AiDotNetTests.csprojcleandotnet build AiDotNetBenchmarkTests/AiDotNetBenchmarkTests.csprojclean🤖 Generated with Claude Code