Skip to content

[Performance] Attention node runs on CPU regardless of provider #28025

Description

@ir2718

Describe the issue

By creating a model using the Attention node, the node gets placed on CPU regardless of the provider set in the providers argument.

To reproduce

Virtual environment setup:

pip install onnx==1.21.0 onnxruntime-gpu==1.24.4 torch==2.10.0

Script used to create the model:

import onnx

batch_size = 1
seq_len = 8
hidden_dim = 64
num_heads = 8

q = onnx.helper.make_tensor_value_info(
    "q", onnx.TensorProto.FLOAT, [batch_size, seq_len, hidden_dim * num_heads])
k = onnx.helper.make_tensor_value_info(
    "k", onnx.TensorProto.FLOAT, [batch_size, seq_len, hidden_dim * num_heads])
v = onnx.helper.make_tensor_value_info(
    "v", onnx.TensorProto.FLOAT, [batch_size, seq_len, hidden_dim * num_heads])
o = onnx.helper.make_tensor_value_info(
    "o", onnx.TensorProto.FLOAT, [batch_size, seq_len, hidden_dim * num_heads])
attention_node = onnx.helper.make_node(
    "Attention",
    inputs=["q", "k", "v"],
    outputs=["o"],
    q_num_heads=num_heads,
    kv_num_heads=num_heads,
    domain=""
)
graph = onnx.helper.make_graph(
    nodes=[attention_node],
    name="attention_graph",
    inputs=[q, k, v],
    outputs=[o],
)
model = onnx.helper.make_model(
    graph,
    opset_imports=[onnx.helper.make_opsetid("", 24)],
    producer_name="attention-example",
)
onnx.checker.check_model(model)
onnx.save_model(model, "model.onnx")

The script used for running the model:

import torch
import onnxruntime as ort

batch_size = 1
seq_len = 8
hidden_dim = 64
num_heads = 8

opts = ort.SessionOptions()
opts.enable_profiling = True
ort_model = ort.InferenceSession(
    "model.onnx", opts,
    providers=["CUDAExecutionProvider"],
    provider_options=[{
        "device_id": 0,
    }]
)
init_tensor_fn = lambda: torch.randn(
    batch_size, seq_len, hidden_dim * num_heads)
q, k, v = init_tensor_fn(), init_tensor_fn(), init_tensor_fn(),
out = ort_model.run(None, {
    "q": q.numpy(),
    "k": k.numpy(),
    "v": v.numpy(),
})

The generated profiling output:

[
{"cat" : "Session","pid" :1639450,"tid" :1639450,"dur" :250,"ts" :4,"ph" : "X","name" :"model_loading_uri","args" : {}},
{"cat" : "Session","pid" :1639450,"tid" :1639450,"dur" :13824,"ts" :113822,"ph" : "X","name" :"session_initialization","args" : {}},
{"cat" : "Node","pid" :1639450,"tid" :1639450,"dur" :69,"ts" :128763,"ph" : "X","name" :"Attention_0_kernel_time","args" : {"thread_scheduling_stats" : {"main_thread": {"thread_pool_name": "session-1-intra-op", "thread_id": "139802086621440", "block_size": [], "core": -1, "Distribution": 0, "DistributionEnqueue": 0, "Run": 0, "Wait": 0, "WaitRevoke": 0}, "sub_threads": {"139795150927552": {"num_run": 0, "core": -1},"139795142534848": {"num_run": 0, "core": -1},"139795134142144": {"num_run": 0, "core": -1},"139795125749440": {"num_run": 0, "core": -1},"139795117356736": {"num_run": 0, "core": -1},"139795108964032": {"num_run": 0, "core": -1},"139795100571328": {"num_run": 0, "core": -1},"139795092178624": {"num_run": 0, "core": -1},"139795011466944": {"num_run": 0, "core": -1},"139795003074240": {"num_run": 0, "core": -1},"139794994681536": {"num_run": 0, "core": -1},"139794986288832": {"num_run": 0, "core": -1},"139794841597632": {"num_run": 0, "core": -1},"139794977896128": {"num_run": 0, "core": -1},"139794969503424": {"num_run": 0, "core": -1},"139794961110720": {"num_run": 0, "core": -1},"139794877249216": {"num_run": 0, "core": -1},"139794868856512": {"num_run": 0, "core": -1},"139794860463808": {"num_run": 0, "core": -1},"139794852071104": {"num_run": 0, "core": -1},"139794833204928": {"num_run": 0, "core": -1},"139794824812224": {"num_run": 0, "core": -1},"139793736398528": {"num_run": 0, "core": -1}}},"output_type_shape" : [{"float":[1,8,512]}],"output_size" : "16384","parameter_size" : "0","activation_size" : "49152","node_index" : "0","input_type_shape" : [{"float":[1,8,512]},{"float":[1,8,512]},{"float":[1,8,512]}],"provider" : "CPUExecutionProvider","op_name" : "Attention"}},
{"cat" : "Session","pid" :1639450,"tid" :1639450,"dur" :80,"ts" :128759,"ph" : "X","name" :"SequentialExecutor::Execute","args" : {}},
{"cat" : "Session","pid" :1639450,"tid" :1639450,"dur" :106,"ts" :128749,"ph" : "X","name" :"model_run","args" : {}}
]

Urgency

It could be urgent as this is often the slowest operation operation in a model, and it affects all transformer models where an Attention node is used.

Platform

Linux

OS Version

Debian GNU/Linux 13

ONNX Runtime Installation

Released Package

ONNX Runtime Version or Commit ID

1.24.4

ONNX Runtime API

Python

Architecture

X86

Execution Provider

CUDA

Execution Provider Library Version

12.4

Model File

This is the same model generated by the script mentioned in the steps to reproduce.

model.zip

Is this a quantized model?

No

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    performanceissues related to performance regressions

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions