Skip to content

cudaErrorInvalidConfiguration in FusedBatchNormV3 #7316

Description

@bdnkth

Describe the bug
I have custom trained the EfficientDetD0 model from the TensorFlow model zoo for object detection and exported the model to onnx with tf2onnx using opset 11 and fixed input of [1,512,512,3].
Using that onnx model in the onnx runtime for C++ I run into a cudaErrorInvalidConfiguration in FusedBatchNormV3_528 of the EfficientDet.
Full errr message is:
2021-04-12 14:58:25.6606830 [E:onnxruntime:, sequential_executor.cc:339 onnxruntime::SequentialExecutor::Execute] Non-zero status code returned while running Transpose node. Name:'StatefulPartitionedCall/EfficientDet-D0/bifpn/node_03/1_dn_lvl_5/input_0_up_lvl_5/1x1_pre_sample/batchnorm/FusedBatchNormV3__528' Status Message: CUDA error cudaErrorInvalidConfiguration:invalid configuration argument

Our old onnx models which were made in TF 1.15 are running without an error through the code. TF 2.4.0 models do not.

System information

  • OS Platform and Distribution (e.g., Linux Ubuntu 16.04): Windows 10
  • ONNX Runtime installed from (source or binary): binary
  • ONNX Runtime version: 1.7.0
  • Python version: 2.4.1
  • Visual Studio version (if applicable): 16.9.3
  • GCC/Compiler version (if compiling from source): v142
  • CUDA/cuDNN version: 11.0
  • GPU model and memory: RTX 3090 24gb

To Reproduce

  • The EfficientDetD0 onnx model is attached: model.zip

Expected behavior
There should be no cudaErrorInvalidConfiguration

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

ep:CUDAissues related to the CUDA execution provider

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions