Skip to content

rocblas alt impl during backward pass only - #13352

Merged
cloudhan merged 6 commits into
microsoft:mainfrom
ROCm:rocblas-alt-impl-v2
Nov 9, 2022
Merged

cloudhan merged 6 commits into
microsoft:mainfrom
ROCm:rocblas-alt-impl-v2

Conversation

@jeffdaily

@jeffdaily Jeff Daily (jeffdaily) commented Oct 18, 2022 •

Copy link
Copy Markdown
Contributor

On AMD Instinct MI200 GPUs, the FP16 and BF16 V_DOT2 and MFMA matrix instructions flush input and output denormal values to zero. When training using FP16 precision, some models may fail to converge with FP16 denorms flushed to zero. The affected instructions are only used by rocBLAS (GEMM) and MIOpen (convolution) kernels; all other onnxruntime operations will not encounter this behavior. All other supported AMD GPUs will not encounter this behavior.

rocBLAS and MIOpen provide alternate implementations for affected FP16 operations. Alternate implementations for BF16 operations are not provided; BF16 numbers have a larger dynamic range than FP16 numbers and are less likely to encounter denormal values. For the FP16 alternate implementations, FP16 input values are cast to an intermediate BF16 value and then cast back to FP16 output after the accumulate FP32 operations. In this way, the input and output types are unchanged.

Denormal values more frequently occur in the backward pass of training during gradient calculation. Therefore, it is necessary to track when the backward pass of training is executing. For the ROCm EP only, the __backwardpass attribute is added to all Nodes after the YieldOp is detected. This takes place in a level1 graph optimization pass. The attribute is forwarded to any newly created FusedMatMul Nodes. In addition, the scope-based helper class BackwardPassGuard is provided to toggle state for rocblas. This behavior of using the alternate implementations during the backward pass is made automatic with this PR. This default behavior can be overridden using environment variables, ROCBLAS_INTERNAL_FP16_ALT_IMPL and MIOPEN_DEBUG_CONVOLUTION_ATTRIB_FP16_ALT_IMPL. The behavior of these environment variables is as follows:

forward backward
Env unset original alternate
Env set to 1 alternate alternate
Env set to 0 original original

See also:

https://pytorch.org/docs/stable/notes/numerical_accuracy.html#reduced-precision-fp16-and-bf16-gemms-and-convolutions-on-amd-instinct-mi200-devices

@cloudhan

This comment was marked as outdated.

@azure-pipelines

This comment was marked as outdated.

@cloudhan

This comment was marked as outdated.

@azure-pipelines

This comment was marked as outdated.

@cloudhan

This comment was marked as outdated.

@azure-pipelines

This comment was marked as outdated.

@cloudhan

Copy link
Copy Markdown
Contributor

I am wondering whether YieldOp will be rewritten as other op in any cases, say, we have additional transformer passes and YieldOp is written in it. How can we guarantee that the rewritten will correctly propagate the __backwardpass attribute. I'd assume it impose cognitive load if the __backwardpass attribute invariant must be maintained. Shall we add some tests for it?

@jeffdaily

Copy link
Copy Markdown
Contributor Author

I am wondering whether YieldOp will be rewritten as other op in any cases, say, we have additional transformer passes and YieldOp is written in it. How can we guarantee that the rewritten will correctly propagate the __backwardpass attribute. I'd assume it impose cognitive load if the __backwardpass attribute invariant must be maintained. Shall we add some tests for it?

The alternative to forwarding the __backwardpass attribute during transformations would be to run the rocm alt impl transformer as part of each level. I tried this initially, putting the transformation into each opt level's set, but this failed because it can only be registered one time. By design you can elect which graph opt level is used and we need this transformation to be applied as if it were a level 0, and applied after all other optimizations have been applied.

@cloudhan

This comment was marked as outdated.

@azure-pipelines

This comment was marked as outdated.

@cloudhan

Copy link
Copy Markdown
Contributor

Weixing Zhang (@weixingzhang) Could you take a look at this? I do not have the context in #9821

@jeffdaily

Copy link
Copy Markdown
Contributor Author

The PR description is updated. Weixing Zhang (@weixingzhang) please review; this PR is in response to your review of #9821 and we would appreciate your review of the current PR attempt.

@cloudhan
cloudhan merged commit d5d6924 into microsoft:main Nov 9, 2022
MS (simon-moo) pushed a commit to simon-moo/onnxruntime that referenced this pull request Dec 21, 2022
On AMD Instinct MI200 GPUs, the FP16 and BF16 V_DOT2 and MFMA matrix
instructions flush input and output denormal values to zero. When
training using FP16 precision, some models may fail to converge with
FP16 denorms flushed to zero. The affected instructions are only used by
rocBLAS (GEMM) and MIOpen (convolution) kernels; all other onnxruntime
operations will not encounter this behavior. All other supported AMD
GPUs will not encounter this behavior.

rocBLAS and MIOpen provide alternate implementations for affected FP16
operations. Alternate implementations for BF16 operations are not
provided; BF16 numbers have a larger dynamic range than FP16 numbers and
are less likely to encounter denormal values. For the FP16 alternate
implementations, FP16 input values are cast to an intermediate BF16
value and then cast back to FP16 output after the accumulate FP32
operations. In this way, the input and output types are unchanged.

Denormal values more frequently occur in the backward pass of training
during gradient calculation. Therefore, it is necessary to track when
the backward pass of training is executing. For the ROCm EP only, the
`__backwardpass` attribute is added to all Nodes after the YieldOp is
detected. This takes place in a level1 graph optimization pass. The
attribute is forwarded to any newly created FusedMatMul Nodes. In
addition, the scope-based helper class `BackwardPassGuard` is provided
to toggle state for rocblas. This behavior of using the alternate
implementations during the backward pass is made automatic with this PR.
This default behavior can be overridden using environment variables,
ROCBLAS_INTERNAL_FP16_ALT_IMPL and
MIOPEN_DEBUG_CONVOLUTION_ATTRIB_FP16_ALT_IMPL. The behavior of these
environment variables is as follows:

|              | forward   | backward  |
|--------------|-----------|-----------|
| Env unset    | original  | alternate |
| Env set to 1 | alternate | alternate |
| Env set to 0 | original  | original  |

See also:


https://pytorch.org/docs/stable/notes/numerical_accuracy.html#reduced-precision-fp16-and-bf16-gemms-and-convolutions-on-amd-instinct-mi200-devices
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants