xe: fix strided destination handling - #5691
Open
rjoursler wants to merge 2 commits into
Open
Conversation
Add a helper function to simplify size calculations for memory buffers with a different types, but the same layout.
Several GPU implementations operate on temporary buffers with the same layout as the destination. These implementations relied on destination element count for memory size calculations resulting in a GPU segfaults for strided destinations. For testing, the matmul implementation will be covered by the upcoming generated suite. This patch does not add any regression tests for the reference version as the optimized implementations already pass.
Contributor
Author
|
make test |
Contributor
|
While this fixes the segfaults, subbyte_pack now overwrites the strided region... |
Contributor
Author
Agreed, it's not really clear to me what our requirements are in regards to this. I can't say I have ever seen a use case where there are partial updates to a buffer applied via strides, and we aren't exactly testing that these regions are maintained correctly, hence going with the "easy" solution. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes MFDNN-15352.
Several GPU implementations operate on temporary buffers with the same layout as
the destination. These implementations relied on destination element count for
memory size calculations resulting in a GPU segfaults for strided destinations.
For testing, the matmul implementation will be covered by the upcoming generated
suite. This patch does not add any regression tests for the reference version as
the optimized implementations already pass.