Improve cache locality and perf of DeepGru on CPU - #13582
Merged
Merged
Conversation
added 2 commits
November 7, 2022 14:48
Dmitri Smirnov (yuslepukhin)
requested review from
Ashwini Khade (askhade) and
Yufeng Li (yufenglee)
November 7, 2022 22:50
MS (simon-moo)
pushed a commit
to simon-moo/onnxruntime
that referenced
this pull request
Dec 21, 2022
### Description <!-- Describe your changes. --> Introduce Gemm weights pre-pack. ### Motivation and Context A 1-P customer requested a performance improvement for DeepGru which consumes a bulk of CPU in their model. This provides measurable performance improvements. Customer model numbers. gru: mean = 356 us; 1ms = 99.8 prctile; 99th prctile = 665 ms (yuslepukhin/deep_gru_opt) main: mean = 375 us; 1ms = 99.8 prctile; 99th prctile = 695 ms (where yuslepukhin/deep_gru_opt branched off main) 1.13.1: mean = 391 us; 1ms = 99.6 prctile; 99th prctile = 744 ms
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Introduce Gemm weights pre-pack.
Motivation and Context
A 1-P customer requested a performance improvement for DeepGru which consumes a bulk of CPU in their model. This provides measurable performance improvements.
Customer model numbers.
gru: mean = 356 us; 1ms = 99.8 prctile; 99th prctile = 665 ms (yuslepukhin/deep_gru_opt)
main: mean = 375 us; 1ms = 99.8 prctile; 99th prctile = 695 ms (where yuslepukhin/deep_gru_opt branched off main)
1.13.1: mean = 391 us; 1ms = 99.6 prctile; 99th prctile = 744 ms