Summary
Two forecasting models pass their clone-parity (output-correctness) tests after the fixes on fix/ci-baseline-bugs (PR #1455), but still fail their model-family tests due to performance, not divergence:
RWKVForecasterTests.LossStrictlyDecreasesOnMemorizationTask — exceeds the 180s xUnit budget. The memorization task runs many train steps; each RWKV forward+backward over the sequence is dominated by RWKV7Block per-step throughput. (Passes the loss-decrease check when it completes; it just doesn't complete in time.)
MGTSDTests.Clone_ShouldProduceIdenticalOutput — times out at 120s. The clone test runs Predict twice (original + clone); MGTSD's reverse-diffusion sampling loop (_diffusionSteps denoising iterations) is too slow per call.
Why this is separate from the clone-parity work
The clone-divergence root causes (stale cached layer refs after deserialize; DeepCopy skipping lazy layers; the seqLen × modelDim flatten anti-pattern) are fixed in PR #1455. These two are purely throughput — the math is correct, the work just doesn't fit the CI timeout.
Suggested approach
Per the project rule: profile with PerfView and fix the bottleneck — do not shrink the model, skip the test, or extend the timeout.
- RWKV7Block: profile the forward/backward; likely scalar inner loops or per-step allocations (cf. the Mamba2
tensor[new[]{...}] multi-dim-indexer allocation bug already fixed — c45c75093). Consider a fused recurrent path similar to CpuEngine.LstmSequenceForward.
- MGTSD: profile the diffusion sampler; reduce per-step overhead (engine FFT/op wrapping, redundant allocations) rather than the step count.
Acceptance
RWKVForecasterTests.LossStrictlyDecreasesOnMemorizationTask and MGTSDTests.Clone_ShouldProduceIdenticalOutput pass within the standard CI budget with no model-size/iteration/timeout changes.
Branch context: fix/ci-baseline-bugs / PR #1455.
Summary
Two forecasting models pass their clone-parity (output-correctness) tests after the fixes on
fix/ci-baseline-bugs(PR #1455), but still fail their model-family tests due to performance, not divergence:RWKVForecasterTests.LossStrictlyDecreasesOnMemorizationTask— exceeds the 180s xUnit budget. The memorization task runs many train steps; each RWKV forward+backward over the sequence is dominated byRWKV7Blockper-step throughput. (Passes the loss-decrease check when it completes; it just doesn't complete in time.)MGTSDTests.Clone_ShouldProduceIdenticalOutput— times out at 120s. The clone test runsPredicttwice (original + clone); MGTSD's reverse-diffusion sampling loop (_diffusionStepsdenoising iterations) is too slow per call.Why this is separate from the clone-parity work
The clone-divergence root causes (stale cached layer refs after deserialize;
DeepCopyskipping lazy layers; theseqLen × modelDimflatten anti-pattern) are fixed in PR #1455. These two are purely throughput — the math is correct, the work just doesn't fit the CI timeout.Suggested approach
Per the project rule: profile with PerfView and fix the bottleneck — do not shrink the model, skip the test, or extend the timeout.
tensor[new[]{...}]multi-dim-indexer allocation bug already fixed —c45c75093). Consider a fused recurrent path similar toCpuEngine.LstmSequenceForward.Acceptance
RWKVForecasterTests.LossStrictlyDecreasesOnMemorizationTaskandMGTSDTests.Clone_ShouldProduceIdenticalOutputpass within the standard CI budget with no model-size/iteration/timeout changes.Branch context:
fix/ci-baseline-bugs/ PR #1455.