Skip to content

[CUDA] QMoE support shared experts - #29028

Draft
Tianlei Wu (tianleiwu) wants to merge 7 commits into
mainfrom
tlwu/qmoe_shared_experts
Draft

Tianlei Wu (tianleiwu) wants to merge 7 commits into
mainfrom
tlwu/qmoe_shared_experts

Conversation

@tianleiwu

@tianleiwu Tianlei Wu (tianleiwu) commented Jun 12, 2026 •

Copy link
Copy Markdown
Contributor

Description

Extend QMoE to support Qwen 3.5/3.6 style shared experts.

Limitation: The intermediate hidden size shall be same among all experts.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR extends the CUDA MoE/QMoE contrib ops to support Qwen3-style always-on shared experts via a new num_shared_experts attribute, where the last num_shared_experts expert slots are always selected and weighted by a per-token sigmoid gate (excluded from routed softmax/top-k/normalization).

Changes:

  • Adds num_shared_experts to the MoE and QMoE operator schemas and documents the fused shared-expert semantics.
  • Updates CUDA routing (Softmax+TopK) kernels and CUDA MoE/QMoE execution to size per-token selection buffers for k_total = k + num_shared_experts and append shared experts with sigmoid-gated weights.
  • Adds CUDA parity tests validating fused shared-expert behavior for both fp32 MoE and INT4-fp16 QMoE.

Reviewed changes

Copilot reviewed 8 out of 8 changed files in this pull request and generated 4 comments.

Show a summary per file
File Description
onnxruntime/test/python/transformers/test_moe_shared_expert_cuda.py New parity tests comparing fused shared-expert outputs vs routed-only + isolated shared-expert reference composition.
onnxruntime/core/graph/contrib_ops/contrib_defs.cc Adds num_shared_experts attribute (with semantics) to MoE and QMoE schemas.
onnxruntime/contrib_ops/cuda/moe/qmoe_kernels.h Extends LaunchSoftmaxTopK API to accept num_shared_experts.
onnxruntime/contrib_ops/cuda/moe/qmoe_kernels.cu Implements shared-expert routing behavior by sorting only routed experts and appending shared experts with sigmoid gates.
onnxruntime/contrib_ops/cuda/moe/moe.cc Updates CUDA MoE to allocate/operate with k_total selection slots and pass num_shared_experts into routing.
onnxruntime/contrib_ops/cuda/moe/moe_quantization.cc Updates CUDA QMoE similarly to use k_total and pass num_shared_experts into routing.
onnxruntime/contrib_ops/cuda/moe/moe_base.h Parses/validates num_shared_experts (and disallows with use_sparse_mixer).
docs/contrib_ops/cuda/moe_qmoe.md Documents num_shared_experts and adds a dedicated “Shared-Expert Fusion” section and test reference.

Comment thread onnxruntime/test/python/transformers/test_moe_shared_expert_cuda.py Outdated
Comment thread onnxruntime/core/graph/contrib_ops/contrib_defs.cc
Comment thread onnxruntime/core/graph/contrib_ops/contrib_defs.cc
Comment thread docs/contrib_ops/cuda/moe_qmoe.md Outdated

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants