Skip to content

[Feature Request] Extend CUDA ONNX Ops to latest opset version #27729

Description

@tianleiwu

Describe the feature request

Summary

  • Strict registry-level audit found 47 ONNX-domain CUDA ops whose highest registered opset is below the latest ONNX changelog opset.
  • Priority breakdown: P1 28, P2 11, P3 8.
  • Method: compare the highest opset registered in onnxruntime/core/providers/cuda/cuda_execution_provider.cc against the latest version entry in build/cuda/Debug/_deps/onnx-src/docs/Changelog.md.
  • This report is intentionally strict about registration coverage. It does not assume a newer schema is automatically backward-compatible with an older kernel registration.

Priority Legend

  • P1: foundational graph/model-import gaps that are more likely to block modern models or common exporter output.
  • P2: secondary math/activation gaps that are real registration misses but are typically less central.
  • P3: deprecated, generator, or otherwise lower-urgency gaps.

Gaps

Priority Operator CUDA Max ONNX Latest Gap Registration Changelog PR
P1 GlobalAveragePool 1 22 21 cc:1843 md:26711 #27733
P1 GlobalMaxPool 1 22 21 cc:1852 md:26780 #27733
P1 ConstantOfShape 9 25 16 cc:1963 md:31634 #27728
P1 TopK 11 24 13 cc:2068 md:31274 #27735
P1 Flatten 13 25 12 cc:2334 md:31731 #27728
P1 RoiAlign 10 22 12 cc:2025 md:27917 #27646
P1 Size 13 25 12 cc:2263 md:32532 #27728
P1 ConvTranspose 11 22 11 cc:2080 md:26246 #27710
P1 MaxPool 12 22 10 cc:2121 md:27215 #27715
P1 Dropout 13 22 9 cc:2326 md:26459 Skip since it is training ops
P1 GRU 14 22 8 cc:2437 md:26601 #27738
P1 LSTM 14 22 8 cc:2440 md:26988 #27737
P1 RNN 14 22 8 cc:2434 md:27625 #27743
P1 Pad 18 25 7 cc:2522 md:32016 #27708
P1 Identity 19 25 6 cc:2567 md:31769 #27728
P1 If 19 25 6 cc:2568 md:31798 #27728
P1 Loop 19 25 6 cc:2569 md:31838 #27728
P1 Scan 19 25 6 cc:2590 md:32289 #27728
P1 DequantizeLinear 21 25 4 cc:2603 md:31672 #27745
P1 QuantizeLinear 21 25 4 cc:2617 md:32159 #27741
P1 Cast 23 25 2 cc:2686 md:31420 #27744
P1 Reshape 23 25 2 cc:2707 md:32239 #27742
P1 Shape 23 25 2 cc:2706 md:32455 #27734
P1 Squeeze 23 25 2 cc:2709 md:32563 #27739
P1 Transpose 23 25 2 cc:2708 md:32597 #27740
P1 Unsqueeze 23 25 2 cc:2710 md:32642 #27739
P2 Elu 6 22 16 cc:1677 md:26516 #27753
P2 InstanceNormalization 6 22 16 cc:1937 md:26943 #27758
P2 Selu 6 22 16 cc:1689 md:28021 #27753
P2 Cos 7 22 15 cc:1656 md:26313 #27756
P2 Sin 7 22 15 cc:1659 md:28062 #27756
P2 EyeLike 9 22 13 cc:1985 md:26555 #27757
P2 ThresholdedRelu 10 22 12 cc:2029 md:28209 #27753
P2 Round 11 22 11 cc:2106 md:27979 #27754
P2 Equal 13 19 6 cc:2220 md:22998 #27754
P2 ReduceMax 18 20 2 cc:2515 md:24331 #27755
P2 ReduceMin 18 20 2 cc:2508 md:24380 #27755
P3 RandomNormal 1 22 21 cc:2005 md:27726
P3 RandomNormalLike 1 22 21 cc:2006 md:27772
P3 RandomUniform 1 22 21 cc:2007 md:27822
P3 RandomUniformLike 1 22 21 cc:2008 md:27867
P3 Softplus 1 22 21 cc:1701 md:28120
P3 Softsign 1 22 21 cc:1695 md:28151
P3 Scatter 10 11 1 cc:1986 md:12972
P3 Upsample 9 10 1 cc:1957 md:10297

Notes

  • Scatter and Upsample are deprecated ONNX operators, so they stay in the report but are ranked lower.
  • Several P1 items are only 1-2 opsets behind, but they remain high priority because they are common graph-building or deployment operators.
  • The earlier one-item Upsample note was too narrow; this wider pass covers the full ONNX-domain CUDA registration surface in the execution provider.

Describe scenario use case

To run model exported with lastest onnx opset version

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    ep:CUDAissues related to the CUDA execution providerfeature requestrequest for unsupported feature or enhancement

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions