Disclaimer: AI tooling was used heavily in the writing this issue. Any code not physically typed out by myself has been carefully reviewed to the best of my ability.
Summary
When respv::Optimizer::run evaluates an OpBranchConditional whose condition resolves to a constant and one of its targets is the back-edge of a structured loop, the optimizer cascades into eliminating the merge block, leaving orphaned branch instructions and producing SPIR-V that fails spirv-val with:
error: Branch must appear in a block
This was reported downstream in Godot as godotengine/godot#118389 — a do { } while (false); in a shader hangs vkCreateGraphicsPipelines on Intel/Iris (Windows 10) and reproduces with --gpu-validation on Linux/NVIDIA.
Reproduction
Shader:
#version 450
layout(location = 0) out float outX;
void main() {
do {
} while (false);
outX = 1.0;
}
Compile and strip debug info:
glslangValidator -V dowhile.frag -o dowhile.spv
spirv-opt --strip-debug dowhile.spv -o dowhile_in.spv
Disassembled input (relevant portion):
%4 = OpFunction %void None %3
%5 = OpLabel
OpBranch %6
%6 = OpLabel
OpLoopMerge %8 %9 None
OpBranch %7
%7 = OpLabel
OpBranch %9
%9 = OpLabel
OpBranchConditional %false %6 %8
%8 = OpLabel
OpStore %14 %float_1
OpReturn
OpFunctionEnd
Run through the optimizer (no spec constants needed — the condition is OpConstantFalse):
respv::Shader shader(spv_data, spv_size, /*pInlineFunctions=*/false);
std::vector<uint8_t> out;
respv::Optimizer::run(shader, nullptr, 0, out);
Output:
%4 = OpFunction %void None %3
%5 = OpLabel
OpBranch %6
OpBranch %8
OpFunctionEnd
spirv-val reports Branch must appear in a block on OpBranch %8.
Root cause
In Shader::process (re-spirv.cpp around L1940), back-edges from a loop's continue block to the loop header are intentionally skipped when building the adjacency list, so the topological sort terminates:
// Make sure this label not pointing back to the loop header while on a loop merge.
if (!loopMergeBlockStack.empty() && (labelId == loopMergeBlockStack.back())) {
continue;
}
This means the loop header's recorded in-degree only counts forward predecessors. In the example above, %6's in-degree is 1 (the OpBranch %6 from %5); the back-edge from %9 contributes 0.
In optimizerEvaluateTerminator (L2788), when the condition is constant false, the code calls:
defaultLabelId = optimizedWords[wordIndex + 3]; // %8 (merge)
optimizerReduceLabelDegree(optimizedWords[wordIndex + 2], rContext); // %6 (header)
But %6's in-degree (1) does not account for the back-edge being pruned — decrementing it drops to 0, which triggers cascade-elimination:
%6 block eliminated. OpLoopMerge's labels %8, %9 pushed to the stack; OpBranch %7's label %7 pushed.
%7 block eliminated. %9 pushed.
%9 block eliminated — but the OpBranchConditional is still in its original form (the patch happens after optimizerReduceLabelDegree returns), so iterating its labels pushes both %6 (already at 0, skipped) and %8 again.
%8's in-degree drops from 2 to 0 and %8 is eliminated — collateral damage.
- The function falls through to patch
OpBranchConditional into OpBranch %8, but %8's label and surrounding block are gone, so the branch is orphaned.
The SPIR-V is invalid; some drivers (Intel/Iris) hang in pipeline creation, others throw validation errors.
Proposed fix
Detect back-edges using the pre/post-order block indices already computed in Shader::process. An edge u → v is a back-edge to ancestor v iff pre[u] > pre[v] && post[u] < post[v]. When the pruned target is a back-edge, skip the terminator optimization entirely — the structured loop must remain intact (an OpLoopMerge requires the continue block to branch back to the header, so we can't just drop the back-edge without also rewriting the loop construct).
if (opCode == SpvOpBranchConditional) {
uint32_t prunedLabelId;
if (operatorResolution.value.u32) {
defaultLabelId = optimizedWords[wordIndex + 2];
prunedLabelId = optimizedWords[wordIndex + 3];
} else {
defaultLabelId = optimizedWords[wordIndex + 3];
prunedLabelId = optimizedWords[wordIndex + 2];
}
bool isBackEdge = false;
const uint32_t branchBlockIndex = rContext.shader.instructions[pInstructionIndex].blockIndex;
const uint32_t prunedInstructionIndex = rContext.shader.results[prunedLabelId].instructionIndex;
if ((branchBlockIndex != UINT32_MAX) && (prunedInstructionIndex != UINT32_MAX)) {
const uint32_t prunedBlockIndex = rContext.shader.instructions[prunedInstructionIndex].blockIndex;
if (prunedBlockIndex != UINT32_MAX) {
const uint32_t branchPre = rContext.shader.blockPreOrderIndices[branchBlockIndex];
const uint32_t prunedPre = rContext.shader.blockPreOrderIndices[prunedBlockIndex];
const uint32_t branchPost = rContext.shader.blockPostOrderIndices[branchBlockIndex];
const uint32_t prunedPost = rContext.shader.blockPostOrderIndices[prunedBlockIndex];
isBackEdge = (branchPre > prunedPre) && (branchPost < prunedPost);
}
}
if (isBackEdge) {
return;
}
optimizerReduceLabelDegree(prunedLabelId, rContext);
// ... existing patching code unchanged
}
The same case for OpSwitch does not apply: SPIR-V structured switches cannot have back-edge labels.
Trade-off: a constant-false do-while stays as-is in the optimized SPIR-V instead of being collapsed. That's the same conservative behavior as having a non-constant condition, and downstream driver compilers fold it themselves. I'd rather give that up than rewrite the loop construct here.
I verified locally my my macbook that the patched output passes spirv-val on the reproducer above.
Disclaimer: AI tooling was used heavily in the writing this issue. Any code not physically typed out by myself has been carefully reviewed to the best of my ability.
Summary
When
respv::Optimizer::runevaluates anOpBranchConditionalwhose condition resolves to a constant and one of its targets is the back-edge of a structured loop, the optimizer cascades into eliminating the merge block, leaving orphaned branch instructions and producing SPIR-V that failsspirv-valwith:This was reported downstream in Godot as godotengine/godot#118389 — a
do { } while (false);in a shader hangsvkCreateGraphicsPipelineson Intel/Iris (Windows 10) and reproduces with--gpu-validationon Linux/NVIDIA.Reproduction
Shader:
Compile and strip debug info:
Disassembled input (relevant portion):
Run through the optimizer (no spec constants needed — the condition is
OpConstantFalse):Output:
spirv-valreportsBranch must appear in a blockonOpBranch %8.Root cause
In
Shader::process(re-spirv.cpp around L1940), back-edges from a loop's continue block to the loop header are intentionally skipped when building the adjacency list, so the topological sort terminates:This means the loop header's recorded in-degree only counts forward predecessors. In the example above,
%6's in-degree is1(theOpBranch %6from%5); the back-edge from%9contributes0.In
optimizerEvaluateTerminator(L2788), when the condition is constantfalse, the code calls:But
%6's in-degree (1) does not account for the back-edge being pruned — decrementing it drops to0, which triggers cascade-elimination:%6block eliminated.OpLoopMerge's labels%8,%9pushed to the stack;OpBranch %7's label%7pushed.%7block eliminated.%9pushed.%9block eliminated — but theOpBranchConditionalis still in its original form (the patch happens afteroptimizerReduceLabelDegreereturns), so iterating its labels pushes both%6(already at 0, skipped) and%8again.%8's in-degree drops from2to0and%8is eliminated — collateral damage.OpBranchConditionalintoOpBranch %8, but%8's label and surrounding block are gone, so the branch is orphaned.The SPIR-V is invalid; some drivers (Intel/Iris) hang in pipeline creation, others throw validation errors.
Proposed fix
Detect back-edges using the pre/post-order block indices already computed in
Shader::process. An edgeu → vis a back-edge to ancestorviffpre[u] > pre[v] && post[u] < post[v]. When the pruned target is a back-edge, skip the terminator optimization entirely — the structured loop must remain intact (anOpLoopMergerequires the continue block to branch back to the header, so we can't just drop the back-edge without also rewriting the loop construct).The same case for
OpSwitchdoes not apply: SPIR-V structured switches cannot have back-edge labels.Trade-off: a constant-false do-while stays as-is in the optimized SPIR-V instead of being collapsed. That's the same conservative behavior as having a non-constant condition, and downstream driver compilers fold it themselves. I'd rather give that up than rewrite the loop construct here.
I verified locally my my macbook that the patched output passes
spirv-valon the reproducer above.