Conversation
…, pure predict Three bugs surfaced by the Generated Layers N-Z shard (WorldModelsAgentTests.ActionSelection_ShouldBeFinite / Clone_ShouldProduceSamePolicy): 1. Malformed, multi-GB networks. Each VAE/RNN network was built from an architecture with NO explicit layers, so NeuralNetwork.InitializeLayers auto-generated default hidden layers (~2x the flattened input) and the agent's AddLayer calls then stacked on top. For a 64x64x3 = 12,288-wide observation that default layer alone is a single 12,288x24,576 weight (~2.4 GB at FP64), which tripped weight-streaming's per-tensor byte cap and threw on the first forward. Now each network is built with its EXACT layer list passed to the architecture, so no defaults are generated. 2. Non-deterministic clone. NeuralNetwork.InitializeLayers does not wire explicitly-supplied custom layers, so their lazy weights initialized from the shared non-deterministic RNG — a clone re-initialized a different policy. Each layer now gets a deterministic per-layer RandomSeed, and Clone reproduces the networks via the constructor (they are never trained; only the controller is, via the evolution strategy) and copies the trained controller weights + rollout state directly. A serialization round-trip is NOT used for cloning because it rebuilds the layers and drops the seed pins. 3. Non-pure Predict. SelectAction advances the RNN hidden state for sequential rollout, making repeated Predict calls non-deterministic. Predict now snapshots and restores the hidden state so inference is side-effect-free. All 7 WorldModelsAgentTests pass. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
Contributor
|
Important Review skippedDraft detected. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Pro Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
Owner
Author
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Fixes the Generated Layers N-Z shard from the PR #1563 failing-CI list —
WorldModelsAgentTests.ActionSelection_ShouldBeFiniteandClone_ShouldProduceSamePolicy(in my hand-off lane from #1562; does not touch any files #1562 owns).Three distinct root-cause bugs, found by reproduction:
1. Malformed multi-GB networks → weight-streaming crash
Each VAE/RNN sub-network was constructed from a
NeuralNetworkArchitecturewith no explicit layers, soNeuralNetwork.InitializeLayersfell back toCreateDefaultNeuralNetworkLayers(auto-generating a hidden layer ~2× the flattened input) and the agent'sAddLayercalls then stacked the intended layers on top. For a 64×64×3 = 12,288-wide observation, that default layer alone is a single 12,288×24,576 weight (~2.4 GB at FP64) — which tripped weight-streaming's per-tensorbyte[]cap (int.MaxValue) and threwNotSupportedExceptionon the first forward.Fix: build each network with its exact layer list passed to the architecture, so no default layers are generated.
2. Non-deterministic Clone
InitializeLayersdoes not wire explicitly-supplied custom layers, so their lazy weights initialized from the shared non-deterministic RNG — a clone re-initialized a different policy.Fix: assign a deterministic per-layer
RandomSeed;Clone()reproduces the networks via the constructor (they are never trained — only the controller is, via the evolution strategy) and copies the trained controller weights + rollout state directly. A serialization round-trip is not used for cloning because it rebuilds the layers and drops the seed pins.3. Non-pure Predict
SelectActionadvances the RNN hidden state for sequential rollout, making repeatedPredictcalls non-deterministic (surfacedPolicy_ShouldBeDeterministiconce the crash was gone).Fix:
Predictsnapshots and restores the hidden state so inference is side-effect-free.Verification
All 7
WorldModelsAgentTestspass locally (net10.0, Release):ActionSelection_ShouldBeFinite,Clone_ShouldProduceSamePolicy,Policy_ShouldBeDeterministic, and 4 others.🤖 Generated with Claude Code