Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 3 additions & 1 deletion cmake/onnxruntime_unittests.cmake
Original file line number Diff line number Diff line change
Expand Up @@ -792,7 +792,9 @@ if(onnxruntime_USE_JSEP)
endif()

if(onnxruntime_USE_WEBGPU AND NOT onnxruntime_USE_EP_API_ADAPTERS)
list(APPEND onnxruntime_test_framework_src_patterns ${TEST_SRC_DIR}/providers/webgpu/*)
list(APPEND onnxruntime_test_framework_src_patterns
${TEST_SRC_DIR}/providers/webgpu/*
${TEST_SRC_DIR}/providers/webgpu/math/*)
list(APPEND onnxruntime_test_providers_dependencies onnxruntime_providers_webgpu)
list(APPEND onnxruntime_test_providers_libs onnxruntime_providers_webgpu)
endif()
Expand Down
29 changes: 20 additions & 9 deletions docs/ContribOperators.md
Original file line number Diff line number Diff line change
Expand Up @@ -219,7 +219,6 @@ This version of the operator has been available since version 1 of the 'com.micr
<dd>Constrain mask index to integer types</dd>
</dl>


### <a name="com.microsoft.AttnLSTM"></a><a name="com.microsoft.attnlstm">**com.microsoft.AttnLSTM**</a>

Computes an one-layer RNN where its RNN Cell is an AttentionWrapper wrapped a LSTM Cell. The RNN layer
Expand Down Expand Up @@ -5661,8 +5660,13 @@ This version of the operator has been available since version 1 of the 'com.micr
If block_size is provided, both hidden_size and inter_size must be divisible by the block size, and
the dequantization is performed per block of size block_size along the K (input feature) dimension.

If block_size and zero_point are provided, both hidden_size and inter_size must be divisible by block_size * pack_size,
where pack_size = 8 / expert_weight_bits.
Packed byte dimensions are computed as logical_element_count * effective_expert_weight_bits / 8.
Weight rows must be byte-aligned. Zero-point rows are padded to a whole byte when necessary.

fc1_expert_weight_bits, fc2_expert_weight_bits, and fc3_expert_weight_bits optionally override
expert_weight_bits for the corresponding projection. An omitted override inherits expert_weight_bits.
When SwiGLU is fused, FC3 is stored in FC1 and fc3_expert_weight_bits must be omitted or equal to
fc1_expert_weight_bits after inheritance.

The SwiGLU (Swish-Gated Linear Unit) activation function is like:
g = xW + b
Expand Down Expand Up @@ -5695,6 +5699,12 @@ This version of the operator has been available since version 1 of the 'com.micr
<dd>Size of each quantization block along the K (input feature) dimension. Must be power of two and ≥ 16 (e.g., 16, 32, 64, 128). Both hidden_size and inter_size must be divisible by the block size. The FP4 modes always use blocking: MXFP4 ('fp4'/'wfp4afp8') is normalized to block_size 32 and NVFP4 ('nvfp4') to block_size 16, even when block_size is omitted. For integer quantization ('int'), omitting block_size means there is no blocking and a whole column shares one scaling factor. </dd>
<dt><tt>expert_weight_bits</tt> : int</dt>
<dd>Number of bits used in quantized weights. Supported values are 2, 4, and 8. Default is 4 bits</dd>
<dt><tt>fc1_expert_weight_bits</tt> : int</dt>
<dd>Optional FC1 override for expert_weight_bits. Inherits expert_weight_bits when omitted.</dd>
<dt><tt>fc2_expert_weight_bits</tt> : int</dt>
<dd>Optional FC2 override for expert_weight_bits. Inherits expert_weight_bits when omitted.</dd>
<dt><tt>fc3_expert_weight_bits</tt> : int</dt>
<dd>Optional FC3 override for expert_weight_bits. Inherits expert_weight_bits when omitted. For fused SwiGLU, the effective FC3 width must equal the effective FC1 width.</dd>
<dt><tt>k</tt> : int</dt>
<dd>Number of top experts to select from expert pool</dd>
<dt><tt>normalize_routing_weights</tt> : int</dt>
Expand All @@ -5719,29 +5729,29 @@ This version of the operator has been available since version 1 of the 'com.micr
<dt><tt>router_probs</tt> : T</dt>
<dd>2D tensor with shape (num_tokens, num_experts)</dd>
<dt><tt>fc1_experts_weights</tt> : T1</dt>
<dd>3D tensor with shape (num_experts, fusion_size * inter_size, hidden_size / pack_size), The fusion_size is 2 for fused swiglu, or 1 otherwise. The pack_size is 8 / expert_weight_bits.</dd>
<dd>3D tensor with shape (num_experts, fusion_size * inter_size, hidden_size * effective_fc1_bits / 8). The last dimension must be byte-aligned. The fusion_size is 2 for fused swiglu, or 1 otherwise. effective_fc1_bits is fc1_expert_weight_bits when provided, otherwise expert_weight_bits.</dd>
<dt><tt>fc1_scales</tt> (optional) : T2</dt>
<dd>Optional weight scales. For quant_type='int', this is a 2D tensor with shape (num_experts, fusion_size * inter_size), or a 3D tensor with shape (num_experts, fusion_size * inter_size, hidden_size / block_size) when block_size is provided. For quant_type='fp4' or 'wfp4afp8', this is a float8e8m0 MXFP block-scale tensor with shape (num_experts, fusion_size * inter_size, hidden_size / 32). For quant_type='nvfp4', this is a float8e4m3fn NVFP4 block-scale tensor with shape (num_experts, fusion_size * inter_size, hidden_size / 16). Not used for quant_type='fp8'.</dd>
<dt><tt>fc1_experts_bias</tt> (optional) : T</dt>
<dd>2D optional tensor with shape (num_experts, fusion_size * inter_size)</dd>
<dt><tt>fc2_experts_weights</tt> : T1</dt>
<dd>3D tensor with shape (num_experts, hidden_size, inter_size / pack_size)</dd>
<dd>3D tensor with shape (num_experts, hidden_size, inter_size * effective_fc2_bits / 8). The last dimension must be byte-aligned. effective_fc2_bits is fc2_expert_weight_bits when provided, otherwise expert_weight_bits.</dd>
<dt><tt>fc2_scales</tt> (optional) : T2</dt>
<dd>Optional weight scales. For quant_type='int', this is a 2D tensor with shape (num_experts, hidden_size), or a 3D tensor with shape (num_experts, hidden_size, inter_size / block_size) when block_size is provided. For quant_type='fp4' or 'wfp4afp8', this is a float8e8m0 MXFP block-scale tensor with shape (num_experts, hidden_size, inter_size / 32). For quant_type='nvfp4', this is a float8e4m3fn NVFP4 block-scale tensor with shape (num_experts, hidden_size, inter_size / 16). Not used for quant_type='fp8'.</dd>
<dt><tt>fc2_experts_bias</tt> (optional) : T</dt>
<dd>2D optional tensor with shape (num_experts, hidden_size)</dd>
<dt><tt>fc3_experts_weights</tt> (optional) : T1</dt>
<dd>3D optional tensor with shape (num_experts, inter_size, hidden_size / pack_size)</dd>
<dd>3D optional tensor with shape (num_experts, inter_size, hidden_size * effective_fc3_bits / 8). The last dimension must be byte-aligned. effective_fc3_bits is fc3_expert_weight_bits when provided, otherwise expert_weight_bits.</dd>
<dt><tt>fc3_scales</tt> (optional) : T2</dt>
<dd>Optional weight scales. For quant_type='int', this is a 2D tensor with shape (num_experts, inter_size), or a 3D tensor with shape (num_experts, inter_size, hidden_size / block_size) when block_size is provided. For quant_type='fp4' or 'wfp4afp8', this is a float8e8m0 MXFP block-scale tensor with shape (num_experts, inter_size, hidden_size / 32). Not used for quant_type='fp8'.</dd>
<dt><tt>fc3_experts_bias</tt> (optional) : T</dt>
<dd>2D optional tensor with shape (num_experts, inter_size)</dd>
<dt><tt>fc1_zero_points</tt> (optional) : T1</dt>
<dd>2D tensor with shape (num_experts, fusion_size * inter_size / pack_size), or 3D tensor with shape (num_experts, fusion_size * inter_size, hidden_size / block_size / pack_size) when block_size is provided.</dd>
<dd>2D tensor with shape (num_experts, ceil(fusion_size * inter_size * effective_fc1_bits / 8)), or 3D tensor with shape (num_experts, fusion_size * inter_size, ceil((hidden_size / block_size) * effective_fc1_bits / 8)) when block_size is provided.</dd>
<dt><tt>fc2_zero_points</tt> (optional) : T1</dt>
<dd>2D tensor with shape (num_experts, hidden_size / pack_size), or 3D tensor with shape (num_experts, hidden_size, inter_size / block_size / pack_size) when block_size is provided.</dd>
<dd>2D tensor with shape (num_experts, ceil(hidden_size * effective_fc2_bits / 8)), or 3D tensor with shape (num_experts, hidden_size, ceil((inter_size / block_size) * effective_fc2_bits / 8)) when block_size is provided.</dd>
<dt><tt>fc3_zero_points</tt> (optional) : T1</dt>
<dd>2D optional tensor with shape (num_experts, inter_size / pack_size), or 3D optional tensor with shape (num_experts, inter_size, hidden_size / block_size / pack_size) when block_size is provided.</dd>
<dd>2D optional tensor with shape (num_experts, ceil(inter_size * effective_fc3_bits / 8)), or 3D optional tensor with shape (num_experts, inter_size, ceil((hidden_size / block_size) * effective_fc3_bits / 8)) when block_size is provided.</dd>
<dt><tt>router_weights</tt> (optional) : T</dt>
<dd>2D optional tensor with shape (num_tokens, num_experts). When provided, router_probs is used only for Top-K expert selection, and router_weights is used for aggregating expert outputs (the values at the selected expert indices are gathered and used as mixing weights). This enables DeepSeek-style noaux_tc routing where different tensors are used for selection and aggregation. When not provided, router_probs is used for both selection and aggregation (backward compatible).</dd>
<dt><tt>fc1_global_scale</tt> (optional) : T4</dt>
Expand Down Expand Up @@ -7749,3 +7759,4 @@ No versioning maintained for experimental ops.
<dd>Constrain input and output types to float32 tensors.</dd>
</dl>


11 changes: 11 additions & 0 deletions docs/contrib_ops/cuda/paged_attention.md
Original file line number Diff line number Diff line change
Expand Up @@ -235,6 +235,7 @@ ops without translation.
| `kv_num_heads` | INT | required | existing |
| `scale` | FLOAT | `1/sqrt(head_size)` | existing — mandatory in `LATENT` (§12.6) |
| `softcap` | FLOAT | `0.0` | existing |
| `is_causal` | INT | `1` | `0` removes the right-hand causal bound on all CUDA backends |
| `local_window_size` | INT | `-1` | existing — §9 |
| `do_rotary` | INT | `0` | existing |
| `rotary_interleaved` | INT | `0` | existing |
Expand Down Expand Up @@ -1630,6 +1631,16 @@ These block the feature work and should land ahead of it.
> falls back to the memory-efficient backend, which gathers pages into a dense buffer first and
> therefore accepts any block size. The op only errors when neither backend is eligible.
> Lifting this properly requires teaching the Flash paged loader to split a tile across pages.
>
> **Non-causal attention.** `is_causal=0` works with all CUDA backends: FlashAttention,
> memory-efficient attention, paged decode (including XQA), and latent attention. In particular,
> a native cache with `head_size=128` and 16-, 32-, or 64-token pages uses MEA for prefill/multi-token
> drafting and paged decode for decode-shaped batches, without requiring Flash-compatible pages.
> Each query can attend through the sequence's full live KV length (`past_seqlens + query_length`).
> A positive `local_window_size` still bounds the left side at `query_position - window_size + 1`;
> the right side remains unbounded. XQA's speculative mask admits every live draft
> token when non-causal; its single-token kernel needs no different mask. Other backend eligibility
> constraints, including XQA's page alignment and MEA's lack of attention-sink support, are unchanged.
2. **Out-of-bounds binary search.** The binary search over `cumulative_seqlens_q` in
`ReshapeAndCache` and `GatherAndExpandPagedKVCache` can yield `batch_id == batch_size` when
`token_id >= cumulative_seqlens_q[batch_size]`, producing OOB reads of `past_seqlens` and
Expand Down
Loading
Loading