Avoid ForwardDiff Jacobians for GPU steady-state adjoints - #243
Merged
ChrisRackauckas merged 1 commit intoAug 28, 2026
Merged
ChrisRackauckas merged 1 commit into
ChrisRackauckas merged 1 commit into
Conversation
Select finite-difference Jacobian construction for GPU steady-state adjoints while preserving ForwardDiff on CPU states. Add a regression test for both choices. Co-Authored-By: Chris Rackauckas <accounts@chrisrackauckas.com> Co-Authored-By: Codex <noreply@openai.com> Agent-Harness: Codex CLI 0.150.1 Agent-Model: unknown Agent-Session: 01a0441e-3d2f-7170-ab6e-58ef72975a9f
ChrisRackauckas
marked this pull request as ready for review
August 28, 2026 01:32
This was referenced Aug 28, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changed and why
DeepEquilibriumNetworks overrides SciMLSensitivity's automatic steady-state sense-algorithm choice, but its override left
autodiff=truefor GPU states. ForwardDiff 1.4.4 began zeroing the unused tail of chunked dual arrays with scalar indexing, so constructing the steady-state adjoint Jacobian now fails onCuArray.Select
autodiff=falseonly when the steady-state problem's initial state is anAbstractGPUArray, matching SciMLSensitivity's own GPU-specific choice while preserving ForwardDiff Jacobians for CPU states. GPUArraysCore moves from a test extra to a direct dependency because runtime dispatch now uses its publicAbstractGPUArrayinterface.Regression boundary
The current main failure resolves ForwardDiff 1.4.5 and SciMLSensitivity 7.119.0, then errors in
ForwardDiff._seed_zero_partials!withScalar indexing is disallowed:The previous GPU run resolved ForwardDiff 1.4.3 and SciMLSensitivity 7.116.2; its dense steady-state adjoint cases cleared this error before the job later failed in an unrelated V100 cuDNN convolution:
ForwardDiff 1.4.4 introduced the chunk-tail zeroing path:
There were no repository GPU runs between those dependency resolutions, so the resolved dependency boundary, rather than a DeepEquilibriumNetworks commit, is the narrowest available boundary.
Verification
Failing before the source fix, with the new regression test present:
Passing after the fix:
The following also exited successfully with no output:
GPUArraysCore is MIT-licensed, compatible with this MIT-licensed package.
GPU CI classification
The PR's GPU job completed with 578 passes and 48 errors. None of the remaining errors follows the path changed here:
ForwardDiff.Tag{SciMLBase.UDerivativeWrapper...}while constructingSteadyStateAdjointProblem. The PR job has zeroUDerivativeWrapperoccurrences, and the new default-sense-algorithm regression passes 15/15.ForwardDiff.Tag{SciMLBase.NonlinearFunction...}in the forward solve. They come from the pre-existing test configurationNewtonRaphson(; autodiff=AutoForwardDiff(; chunksize=12)), introduced in a30d8f1, and are independent of the default adjoint selected by this PR.A separate focused repair changes that explicit GPU Newton configuration to
AutoFiniteDiff()in Use finite differences for GPU Newton tests #244; it is intentionally not mixed into this adjoint fix.CUDNN_STATUS_EXECUTION_FAILED_CUDARTfrom NNlib/cuDNN forward convolutions in untouched layer and test code. The clean-base job stopped at the original adjoint error before reaching those later cases; the base and PR jobs nevertheless resolve identical GPUArraysCore, Lux, LuxCUDA, NNlib, CUDA, cuDNN, and JLL versions.GPU job: https://github.com/SciML/DeepEquilibriumNetworks.jl/actions/runs/33104383320/job/98635451731
Not verified locally
This host has no CUDA-capable GPU. The regression test verifies the CPU/GPU dispatch decision using an
AbstractGPUArraytest double; the GPU conclusions above come from the clean-base and PR CI logs. Documentation was not built because this does not change public API or documentation.Related documentation fix
The clean-main
DEQsdocumentation failure is addressed separately in #245.🤖 Generated with Codex CLI 0.150.1 (model: unknown; session: 01a0441e-3d2f-7170-ab6e-58ef72975a9f)