Skip to content

fix(#1311 cluster-3): snap VLM vision-encoder head count to divide visionDim cleanly - #1397

Merged
ooples merged 1 commit into
masterfrom
fix/issue-1311-vlm-cross-attention
May 20, 2026
Merged

ooples merged 1 commit into
masterfrom
fix/issue-1311-vlm-cross-attention

Conversation

@ooples

@ooples ooples commented May 19, 2026 •

Copy link
Copy Markdown
Owner

Summary

Closes the shape-contract root cause of #1311 (PR #1290 CI Cluster 3 — VLM/Multimodal cross-attention embedding-dim mismatch). 23 of 23 SmolVLM shape failures eliminated. The other affected models (EmotiVoice, RainbowDQN) were already fixed by intervening work; Phi3Vision's remaining failures are foundation-scale OOM/timeout (separate class).

Per-model state on current master (pre-fix in this PR)

Model Pass Fail Notes
EmotiVoiceTests 26 1 mel→encoder projection fix shipped previously (commit e61329b May 14); remaining 1 = MoreData timeout, perf-gap
RainbowDQNAgentTests 7 0 fully fixed by intervening work
SmolVLMTests 2 23 ALL 23 share the shape-mismatch error — fixed by this PR
Phi3VisionTests 2 23 17 OOM + 6 timeout, foundation-scale (336×336 ViT-Large + 3B-param decoder), NOT the shape contract

Root cause (SmolVLM)

SmolVLM defaults: `VisionDim=384, NumHeads=9`. At the vision-encoder MHA construction in `CreateDefaultPixelShuffleProjectorLayers` (and 9 other VLM factories):

```csharp
new MultiHeadAttentionLayer(numHeads > 16 ? 16 : numHeads,
(visionDim) / (numHeads > 16 ? 16 : numHeads))
```

C# integer division: `384 / 9 = 42`. Then `MultiHeadAttentionLayer._embeddingDimension = 9 * 42 = 378` (NOT 384). The QKV weight matrices end up sized `[378, 378]`, but `PatchEmbeddingLayer` upstream emits patch tokens at visionDim=384 — so `ForwardInternal` throws at the very first vision MHA call with:

```
System.ArgumentException : Input embedding dimension (384) does not match
weight dimension (378). Query shape: [1, 256, 384], Weights shape: [378, 378]
```

The 9-heads / 384-vision-dim mismatch is paper-faithful (SmolVLM uses SmolLM's 9-head decoder config) but the vision encoder is SigLIP-Large @ 16 heads × 64 head-dim = 1024 vision-dim — different counts per subsystem. AiDotNet's `SmolVLMOptions` collapses both to a single `NumHeads` knob, so the factory reuses the decoder's 9 for the vision MHA where it doesn't divide.

Fix

Add `ChooseDivisibleHeadConfig(embedDim, requestedHeads, maxHeads = 16)` helper in `LayerHelper` that returns `(heads, headDim)` with `heads * headDim == embedDim` exactly — finds the largest `h ≤ min(requestedHeads, maxHeads)` such that `embedDim % h == 0`. For SmolVLM (visionDim=384, numHeads=9): start at 9, 384%9=6≠0, drop to 8, 384%8=0 ✓ → `(8, 48)`. MHA gets `[384, 384]` weights matching the 384-dim input.

Add `CreateVisionMha(visionDim, numHeads, initializationStrategy?)` shim that applies the helper and returns the configured `MultiHeadAttentionLayer`. Replace all 10 inline `new MultiHeadAttentionLayer(numHeads > 16 ? 16 : numHeads, ...)` call sites across the VLM factories.

Snapping heads downward (vs upward / padding embedDim) keeps every other shape in the chain unchanged — FFN, LayerNorm, downstream Dense all keep their visionDim-wide view. The trade-off is the attention pattern uses slightly fewer heads than the upstream model card; that's strictly more local than reshaping the entire residual stream. Per-subsystem head counts on the options class (`NumVisionHeads` vs `NumDecoderHeads`) is the paper-faithful long-term fix but is an API-surface change; the minimal no-surface-change fix is to snap downward.

Affected factories (10 sites all converted)

All `(visionDim) / (numHeads > 16 ? 16 : numHeads)` patterns: `CreateDefaultEncoderDecoderVLMLayers`, `CreateDefaultVisualExpertVLMLayers`, `CreateDefaultCrossAttentionResamplerVLMLayers`, `CreateDefaultPixelShuffleProjectorLayers` (SmolVLM — direct fix here), `CreateDefaultVisionAdapterLayers` (Phi3Vision), `CreateDefaultTokenReductionVLMLayers` (DeepSeek-VL), + 4 more.

Defensive fix: applies to any VLM where the configured `NumHeads` doesn't divide `VisionDim` cleanly — not just SmolVLM's specific 9/384 mismatch.

Verification

Pre-fix on master:

```
$ dotnet test --filter "FullyQualifiedName~SmolVLMTests"
Failed: 23, Passed: 2
All 23 failures: "Input embedding dimension (384) does not match weight dimension (378)"
```

Post-fix:

```
$ dotnet test --filter "FullyQualifiedName~SmolVLMTests"
Failed: 14, Passed: 11
Remaining 14: 7 OutOfMemoryException + 6 timeout 120s + 1 timeout 180s
NO MORE shape mismatch failures.
```

This PR closes 23 of 23 SmolVLM shape-contract failures. Remaining 14 SmolVLM + 23 Phi3Vision failures are foundation-scale resource issues (OOM at ImageNet-scale ViT-Large + multi-billion-param decoders), same class as #1394 (ResNet/VGG ImageNet-scale perf). Different root cause; separate follow-up.

Test plan

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Refactor
    • Improved vision and encoder layer construction so attention head configuration is chosen to evenly match model dimensions, yielding more consistent and reliable layer behavior and initialization.

Review Change Stack

Copilot AI review requested due to automatic review settings May 19, 2026 22:30
@vercel

vercel Bot commented May 19, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
aidotnet_website Ready Ready Preview, Comment May 20, 2026 1:02am
aidotnet-playground-api Ready Ready Preview, Comment May 20, 2026 1:02am

@coderabbitai

coderabbitai Bot commented May 19, 2026 •

Copy link
Copy Markdown
Contributor

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 1b6eea5d-d00e-48ae-bf20-ecc418b14a9c

📥 Commits

Reviewing files that changed from the base of the PR and between c074340 and 8e17931.

📒 Files selected for processing (1)
  • src/Helpers/LayerHelper.cs

Walkthrough

LayerHelper centralizes MultiHeadAttention configuration by adding ChooseDivisibleHeadConfig to compute a heads/headDim pair that exactly divides embedDim and CreateVisionMha to construct MultiHeadAttentionLayer<T> using that pair. Ten vision/encoder factory sites were updated to call CreateVisionMha(...) (one forwards initializationStrategy: lazy).

Changes

MultiHeadAttention Configuration Centralization

Layer / File(s) Summary
Head configuration validation and MHA construction helpers
src/Helpers/LayerHelper.cs
ChooseDivisibleHeadConfig validates embedDim and maxHeads, clamps requestedHeads, then decrements heads until it divides embedDim evenly, falling back to 1. CreateVisionMha constructs MultiHeadAttentionLayer<T> with the computed (heads, headDim) and forwards an optional initializationStrategy.
Vision and encoder layer factory call site updates
src/Helpers/LayerHelper.cs
Ten factory method call sites across vision and encoder sequences replace inline MultiHeadAttentionLayer<T> instantiation with CreateVisionMha(visionDim, numHeads); one call site forwards initializationStrategy: lazy through the helper.

Sequence Diagram

No sequence diagram required—the changes are a single-file refactor without multi-component sequential interactions.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Poem

Heads aligned, dimensions fall in place,
One helper carves the divisible space,
Ten factories now call the same name,
Refactor tight, configuration tame,
Code sings neat in orderly grace.

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately captures the main change: adding logic to snap VLM vision-encoder head counts to divide visionDim cleanly, fixing shape-contract failures.
Docstring Coverage ✅ Passed Docstring coverage is 92.31% which is sufficient. The required threshold is 80.00%.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/issue-1311-vlm-cross-attention

Comment @coderabbitai help to get the list of available commands and usage tips.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot encountered an error and was unable to review this pull request. You can try again by re-requesting a review.

…sionDim cleanly

PR #1290 CI Cluster 3 #1311: 23 SmolVLM tests failing on master with the cluster's signature shape-mismatch:

  System.ArgumentException : Input embedding dimension (384) does not match
  weight dimension (378). Query shape: [1, 256, 384], Weights shape: [378, 378]

## Root cause

SmolVLM defaults: VisionDim=384, NumHeads=9. At the vision-encoder MHA construction in `CreateDefaultPixelShuffleProjectorLayers` (and 9 other VLM factories):

  new MultiHeadAttentionLayer<T>(numHeads > 16 ? 16 : numHeads,
                                 (visionDim) / (numHeads > 16 ? 16 : numHeads))

C# integer division: `384 / 9 = 42`. Then `MultiHeadAttentionLayer._embeddingDimension = 9 * 42 = 378` (NOT 384). The QKV weight matrices end up sized `[378, 378]`, but `PatchEmbeddingLayer` upstream emits patch tokens at visionDim=384 — so `ForwardInternal` throws at the very first vision MHA call.

The 9-heads / 384-vision-dim mismatch is paper-faithful (SmolVLM uses SmolLM's 9-head decoder config) but the vision encoder is SigLIP-Large @ 16 heads × 64 head-dim = 1024 vision-dim — different counts per subsystem. AiDotNet's `SmolVLMOptions` collapses both to a single `NumHeads` knob, so the factory reuses the decoder's 9 for the vision MHA where it doesn't divide.

Per-subsystem head counts on the options class (`NumVisionHeads` vs `NumDecoderHeads`) is the paper-faithful long-term fix but is an API-surface change. The minimal, no-surface-change fix is to snap the vision MHA's head count downward to the largest divisor of visionDim that's ≤ numHeads.

## Fix

Add `ChooseDivisibleHeadConfig(embedDim, requestedHeads, maxHeads = 16)` helper in `LayerHelper<T>` that returns `(heads, headDim)` with `heads * headDim == embedDim` exactly — finds the largest `h ≤ min(requestedHeads, maxHeads)` such that `embedDim % h == 0`. For SmolVLM (visionDim=384, numHeads=9): start at 9, 384%9=6≠0, drop to 8, 384%8=0 ✓ → `(8, 48)`. MHA gets [384, 384] weights matching the 384-dim input.

Add `CreateVisionMha(visionDim, numHeads, initializationStrategy?)` shim that applies the helper and returns the configured `MultiHeadAttentionLayer<T>`. Replace all 10 inline `new MultiHeadAttentionLayer<T>(numHeads > 16 ? 16 : numHeads, ...)` call sites across the VLM factories.

Snapping heads downward (vs upward / padding embedDim) keeps every other shape in the chain unchanged — FFN, LayerNorm, downstream Dense all keep their visionDim-wide view. The trade-off is the attention pattern uses slightly fewer heads than the upstream model card; that's strictly more local than reshaping the entire residual stream.

## Verification

Pre-fix (current master):

  $ dotnet test --filter "FullyQualifiedName~EmotiVoiceTests|FullyQualifiedName~Phi3VisionTests|FullyQualifiedName~SmolVLMTests|FullyQualifiedName~RainbowDQNAgentTests"
  Failed: 47, Passed: 37
    EmotiVoiceTests: pass=26, fail=1 (timeout)
    Phi3VisionTests: pass=2, fail=23  (all OOM/timeout, foundation-scale)
    RainbowDQNAgentTests: pass=7, fail=0
    SmolVLMTests: pass=2, fail=23  (all shape-mismatch — THIS PR)

Post-fix:

  $ dotnet test --filter "FullyQualifiedName~SmolVLMTests"
  Failed: 14, Passed: 11
    Remaining 14 failures: 7 OutOfMemoryException + 6 timeout 120s + 1 timeout 180s
    — NO MORE shape mismatch.

So this PR closes **23 of 23 SmolVLM shape-contract failures**. The remaining 14 SmolVLM failures (plus Phi3Vision's 23) are foundation-scale resource issues — same class as #1394 (ResNet/VGG ImageNet-scale perf). Different root cause, separate follow-up.

## Affected paths (10 sites)

All `(visionDim) / (numHeads > 16 ? 16 : numHeads)` patterns in VLM factories:
- CreateDefaultEncoderDecoderVLMLayers
- CreateDefaultVisualExpertVLMLayers
- CreateDefaultCrossAttentionResamplerVLMLayers
- CreateDefaultPixelShuffleProjectorLayers (SmolVLM — direct fix here)
- CreateDefaultVisionAdapterLayers (Phi3Vision)
- CreateDefaultTokenReductionVLMLayers (DeepSeek-VL)
- + 4 more

Closes #1311 partially (shape-contract root cause for SmolVLM; defensive fix applied to all 10 vision-encoder MHA sites). Foundation-scale resource residue tracked elsewhere.
@ooples
ooples force-pushed the fix/issue-1311-vlm-cross-attention branch from c074340 to 8e17931 Compare May 20, 2026 01:01
@ooples
ooples merged commit dcc0aef into master May 20, 2026
31 of 45 checks passed
@ooples
ooples deleted the fix/issue-1311-vlm-cross-attention branch May 20, 2026 02:33

This branch was successfully deployed

2 active deployments
Preview – aidotnet_website — 8e179312 Deployed May 20, 2026 by vercel[bot]
Preview – aidotnet-playground-api — 8e179312 Deployed May 20, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[PR #1290 CI Cluster 3] VLM/Multimodal cross-attention embedding-dim mismatch (71 tests)

3 participants