Describe the bug
On model deployment, we found the inference result different between cpu and gpu onnxruntime library while the model is same.
System information
- Linux Debian
- ONNX Runtime installed from (source or binary): Python version: onnxruntime-gpu==1.5.1
- CUDA/cuDNN version: cuda 10.2
- GPU model and memory: Tesla T4
To Reproduce
- Describe steps/code to reproduce the behavior.
- Attach the ONNX model to the issue (where applicable) to expedite investigation.
- Here is the example code to reproduce the bug.
import sys
import struct
import numpy as np
import onnxruntime as ort
DATA_ROOT = './'
model_file = DATA_ROOT + '/fair_seq_encoder_k_v.onnx'
inputs = DATA_ROOT + "/fairseq_encoder_input.bin"
cpu_output_path = DATA_ROOT + "/fairseq_encoder_output_cpu.bin";
with open(inputs, 'rb') as f:
data = f.read()
fmt = 'f' * (515 * 80)
inputs = struct.unpack(fmt, data)
inputs = np.array(inputs, dtype=np.float32).reshape((1, 515, 80))
with open(cpu_output_path, 'rb') as f:
data = f.read()
fmt = 'f' * (10*4*128*192)
cpu_output = struct.unpack(fmt, data)
cpu_output = np.array(cpu_output, dtype=np.float32).reshape((-1))
ort_session = ort.InferenceSession(model_file)
gpu_output = ort_session.run(None, {'encoder_feats': inputs,
'encoder_feats_lengths': np.array([515]),
'beam_size': np.array([10])})
gpu_output = gpu_output[0].reshape((-1))
print(f'maximum difference between onnxruntime cpu & gpu'
f' is {np.max(np.abs(gpu_output - cpu_output)):.4f}')
Expected behavior
Theoretically the difference between gpu and cpu is expected to be a very small number, like 1e-7.
Screenshots

Additional context
Here is the onnx model and test code
Describe the bug
On model deployment, we found the inference result different between cpu and gpu onnxruntime library while the model is same.
System information
To Reproduce
Expected behavior
Theoretically the difference between gpu and cpu is expected to be a very small number, like
1e-7.Screenshots

Additional context
Here is the onnx model and test code