Skip to content

Avoid ForwardDiff Jacobians for GPU steady-state adjoints - #243

Merged
ChrisRackauckas merged 1 commit into
SciML:mainfrom
ChrisRackauckas-Claude:fix/gpu-steady-state-autodiff
Aug 28, 2026
Merged

ChrisRackauckas merged 1 commit into
SciML:mainfrom
ChrisRackauckas-Claude:fix/gpu-steady-state-autodiff

Conversation

@ChrisRackauckas-Claude

@ChrisRackauckas-Claude ChrisRackauckas-Claude commented Aug 27, 2026 •

Copy link
Copy Markdown
Member

Please ignore this draft until it has been reviewed by @ChrisRackauckas.

What changed and why

DeepEquilibriumNetworks overrides SciMLSensitivity's automatic steady-state sense-algorithm choice, but its override left autodiff=true for GPU states. ForwardDiff 1.4.4 began zeroing the unused tail of chunked dual arrays with scalar indexing, so constructing the steady-state adjoint Jacobian now fails on CuArray.

Select autodiff=false only when the steady-state problem's initial state is an AbstractGPUArray, matching SciMLSensitivity's own GPU-specific choice while preserving ForwardDiff Jacobians for CPU states. GPUArraysCore moves from a test extra to a direct dependency because runtime dispatch now uses its public AbstractGPUArray interface.

Regression boundary

The current main failure resolves ForwardDiff 1.4.5 and SciMLSensitivity 7.119.0, then errors in ForwardDiff._seed_zero_partials! with Scalar indexing is disallowed:

The previous GPU run resolved ForwardDiff 1.4.3 and SciMLSensitivity 7.116.2; its dense steady-state adjoint cases cleared this error before the job later failed in an unrelated V100 cuDNN convolution:

ForwardDiff 1.4.4 introduced the chunk-tail zeroing path:

There were no repository GPU runs between those dependency resolutions, so the resolved dependency boundary, rather than a DeepEquilibriumNetworks commit, is the narrowest available boundary.

Verification

Failing before the source fix, with the new regression test present:

steady-state sensitivity defaults: Test Failed at test/utils_tests.jl:48
  Expression: DEQs.default_sensealg(gpu_prob) isa SteadyStateAdjoint{0, false}
   Evaluated: SciMLSensitivity.SteadyStateAdjoint{0, true, ...} isa
              SciMLSensitivity.SteadyStateAdjoint{0, false}

Test Summary:                       | Pass  Fail  Total   Time
Utils Tests                         |   14     1     15  59.8s
  steady-state sensitivity defaults |    1     1      2   3.5s
ERROR: Some tests did not pass: 14 passed, 1 failed, 0 errored, 0 broken.

Passing after the fix:

$ GROUP=Core julia +release --project=. -e 'using Pkg; Pkg.resolve(); Pkg.test()'
Test Summary: | Pass  Total   Time
Utils Tests   |   15     15  58.2s
Test Summary: | Pass  Total      Time
Layers Tests  | 1538   1538  26m27.2s
Testing DeepEquilibriumNetworks tests passed
$ GROUP=QA julia +release --project=. -e 'using Pkg; Pkg.test()'
Test Summary: | Pass  Total     Time
QA            |   20     20  2m03.6s
Testing DeepEquilibriumNetworks tests passed

The following also exited successfully with no output:

julia +release -m Runic --inplace src/DeepEquilibriumNetworks.jl src/utils.jl test/utils_tests.jl
typos Project.toml src/DeepEquilibriumNetworks.jl src/utils.jl test/utils_tests.jl
git diff --check

GPUArraysCore is MIT-licensed, compatible with this MIT-licensed package.

GPU CI classification

The PR's GPU job completed with 578 passes and 48 errors. None of the remaining errors follows the path changed here:

  • The clean-base failure uses ForwardDiff.Tag{SciMLBase.UDerivativeWrapper...} while constructing SteadyStateAdjointProblem. The PR job has zero UDerivativeWrapper occurrences, and the new default-sense-algorithm regression passes 15/15.
  • Twelve errors use ForwardDiff.Tag{SciMLBase.NonlinearFunction...} in the forward solve. They come from the pre-existing test configuration NewtonRaphson(; autodiff=AutoForwardDiff(; chunksize=12)), introduced in a30d8f1, and are independent of the default adjoint selected by this PR.
    A separate focused repair changes that explicit GPU Newton configuration to AutoFiniteDiff() in Use finite differences for GPU Newton tests #244; it is intentionally not mixed into this adjoint fix.
  • Thirty-six errors are CUDNN_STATUS_EXECUTION_FAILED_CUDART from NNlib/cuDNN forward convolutions in untouched layer and test code. The clean-base job stopped at the original adjoint error before reaching those later cases; the base and PR jobs nevertheless resolve identical GPUArraysCore, Lux, LuxCUDA, NNlib, CUDA, cuDNN, and JLL versions.

GPU job: https://github.com/SciML/DeepEquilibriumNetworks.jl/actions/runs/33104383320/job/98635451731

Not verified locally

This host has no CUDA-capable GPU. The regression test verifies the CPU/GPU dispatch decision using an AbstractGPUArray test double; the GPU conclusions above come from the clean-base and PR CI logs. Documentation was not built because this does not change public API or documentation.

Related documentation fix

The clean-main DEQs documentation failure is addressed separately in #245.

🤖 Generated with Codex CLI 0.150.1 (model: unknown; session: 01a0441e-3d2f-7170-ab6e-58ef72975a9f)

Select finite-difference Jacobian construction for GPU steady-state adjoints while preserving ForwardDiff on CPU states. Add a regression test for both choices.

Co-Authored-By: Chris Rackauckas <accounts@chrisrackauckas.com>

Co-Authored-By: Codex <noreply@openai.com>

Agent-Harness: Codex CLI 0.150.1

Agent-Model: unknown

Agent-Session: 01a0441e-3d2f-7170-ab6e-58ef72975a9f
@ChrisRackauckas
ChrisRackauckas marked this pull request as ready for review August 28, 2026 01:32
@ChrisRackauckas
ChrisRackauckas merged commit 62eb37c into SciML:main Aug 28, 2026
8 of 10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants