Skip to content

fix(rl): WorldModelsAgent network sizing, deterministic clone, pure predict - #1567

Closed
ooples wants to merge 1 commit into
masterfrom
fix/worldmodels-agent-clone-finite
Closed

ooples wants to merge 1 commit into
masterfrom
fix/worldmodels-agent-clone-finite

Conversation

@ooples

@ooples ooples commented Jun 11, 2026

Copy link
Copy Markdown
Owner

Summary

Fixes the Generated Layers N-Z shard from the PR #1563 failing-CI list — WorldModelsAgentTests.ActionSelection_ShouldBeFinite and Clone_ShouldProduceSamePolicy (in my hand-off lane from #1562; does not touch any files #1562 owns).

Three distinct root-cause bugs, found by reproduction:

1. Malformed multi-GB networks → weight-streaming crash

Each VAE/RNN sub-network was constructed from a NeuralNetworkArchitecture with no explicit layers, so NeuralNetwork.InitializeLayers fell back to CreateDefaultNeuralNetworkLayers (auto-generating a hidden layer ~2× the flattened input) and the agent's AddLayer calls then stacked the intended layers on top. For a 64×64×3 = 12,288-wide observation, that default layer alone is a single 12,288×24,576 weight (~2.4 GB at FP64) — which tripped weight-streaming's per-tensor byte[] cap (int.MaxValue) and threw NotSupportedException on the first forward.
Fix: build each network with its exact layer list passed to the architecture, so no default layers are generated.

2. Non-deterministic Clone

InitializeLayers does not wire explicitly-supplied custom layers, so their lazy weights initialized from the shared non-deterministic RNG — a clone re-initialized a different policy.
Fix: assign a deterministic per-layer RandomSeed; Clone() reproduces the networks via the constructor (they are never trained — only the controller is, via the evolution strategy) and copies the trained controller weights + rollout state directly. A serialization round-trip is not used for cloning because it rebuilds the layers and drops the seed pins.

3. Non-pure Predict

SelectAction advances the RNN hidden state for sequential rollout, making repeated Predict calls non-deterministic (surfaced Policy_ShouldBeDeterministic once the crash was gone).
Fix: Predict snapshots and restores the hidden state so inference is side-effect-free.

Verification

All 7 WorldModelsAgentTests pass locally (net10.0, Release): ActionSelection_ShouldBeFinite, Clone_ShouldProduceSamePolicy, Policy_ShouldBeDeterministic, and 4 others.

🤖 Generated with Claude Code

…, pure predict

Three bugs surfaced by the Generated Layers N-Z shard
(WorldModelsAgentTests.ActionSelection_ShouldBeFinite / Clone_ShouldProduceSamePolicy):

1. Malformed, multi-GB networks. Each VAE/RNN network was built from an
   architecture with NO explicit layers, so NeuralNetwork.InitializeLayers
   auto-generated default hidden layers (~2x the flattened input) and the
   agent's AddLayer calls then stacked on top. For a 64x64x3 = 12,288-wide
   observation that default layer alone is a single 12,288x24,576 weight
   (~2.4 GB at FP64), which tripped weight-streaming's per-tensor byte cap and
   threw on the first forward. Now each network is built with its EXACT layer
   list passed to the architecture, so no defaults are generated.

2. Non-deterministic clone. NeuralNetwork.InitializeLayers does not wire
   explicitly-supplied custom layers, so their lazy weights initialized from
   the shared non-deterministic RNG — a clone re-initialized a different
   policy. Each layer now gets a deterministic per-layer RandomSeed, and Clone
   reproduces the networks via the constructor (they are never trained; only
   the controller is, via the evolution strategy) and copies the trained
   controller weights + rollout state directly. A serialization round-trip is
   NOT used for cloning because it rebuilds the layers and drops the seed pins.

3. Non-pure Predict. SelectAction advances the RNN hidden state for sequential
   rollout, making repeated Predict calls non-deterministic. Predict now
   snapshots and restores the hidden state so inference is side-effect-free.

All 7 WorldModelsAgentTests pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@vercel

vercel Bot commented Jun 11, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
aidotnet-playground-api Ready Ready Preview, Comment Jun 11, 2026 3:04pm

@coderabbitai

coderabbitai Bot commented Jun 11, 2026

Copy link
Copy Markdown
Contributor

Important

Review skipped

Draft detected.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 9254acc3-f590-482a-bafe-a2697fb06059

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/worldmodels-agent-clone-finite

Comment @coderabbitai help to get the list of available commands and usage tips.

@ooples

ooples commented Jun 11, 2026

Copy link
Copy Markdown
Owner Author

Consolidated into #1565 (single PR for the PR #1563 CI-failure lane, per request). Branch retained but superseded.

@ooples ooples closed this Jun 11, 2026

This branch was successfully deployed

1 active deployment
Preview – aidotnet-playground-api — 3e5c4a5f Deployed Jun 11, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant