[DO NOT MERGE]Qualcomm AI Engine Direct - Mimi Decoder low SQNR reproduce - #1
Draft
winskuo-quic wants to merge 1 commit into
Draft
Conversation
chenweng-quic
pushed a commit
that referenced
this pull request
Jun 9, 2025
Differential Revision: D75911655 Pull Request resolved: pytorch#11344
shewu-quic
pushed a commit
that referenced
this pull request
Aug 1, 2025
BNNS copy crashes the process when the dtypes differ (pytorch#11714). With the example in this PR (pytorch#11714), we crash the process on main. Here is the stack trace from LLDB: ``` Process 19234 stopped * thread #1, queue = 'com.apple.main-thread', stop reason = signal SIGABRT frame #0: 0x0000000190ac9388 libsystem_kernel.dylib`__pthread_kill + 8 libsystem_kernel.dylib`__pthread_kill: -> 0x190ac9388 <+8>: b.lo 0x190ac93a8 ; <+40> 0x190ac938c <+12>: pacibsp 0x190ac9390 <+16>: stp x29, x30, [sp, #-0x10]! 0x190ac9394 <+20>: mov x29, sp (lldb) bt * thread #1, queue = 'com.apple.main-thread', stop reason = signal SIGABRT * frame #0: 0x0000000190ac9388 libsystem_kernel.dylib`__pthread_kill + 8 frame #1: 0x0000000190b0288c libsystem_pthread.dylib`pthread_kill + 296 frame #2: 0x0000000190a0bc60 libsystem_c.dylib`abort + 124 frame pytorch#3: 0x0000000190910174 libsystem_malloc.dylib`malloc_vreport + 892 frame pytorch#4: 0x0000000190913c90 libsystem_malloc.dylib`malloc_report + 64 frame pytorch#5: 0x000000019091821c libsystem_malloc.dylib`___BUG_IN_CLIENT_OF_LIBMALLOC_POINTER_BEING_FREED_WAS_NOT_ALLOCATED + 32 frame pytorch#6: 0x000000019d2f4084 libBNNS.dylib`___lldb_unnamed_symbol1620 + 564 frame pytorch#7: 0x000000019d2f5bac libBNNS.dylib`___lldb_unnamed_symbol1628 + 680 frame pytorch#8: 0x000000019d69ce48 libBNNS.dylib`BNNSCopy + 616 frame pytorch#9: 0x000000030c74d950 _portable_lib.cpython-310-darwin.so`(anonymous namespace)::copy_using_bnns(executorchcoreml::MultiArray const&, executorchcoreml::MultiArray&) + 188 frame pytorch#10: 0x000000030c74cfdc _portable_lib.cpython-310-darwin.so`(anonymous namespace)::copy(executorchcoreml::MultiArray const&, executorchcoreml::MultiArray&, executorchcoreml::MultiArray::CopyOptions) + 72 frame pytorch#11: 0x000000030c74ceec _portable_lib.cpython-310-darwin.so`executorchcoreml::MultiArray::copy(executorchcoreml::MultiArray&, executorchcoreml::MultiArray::CopyOptions) const + 148 frame pytorch#12: 0x000000030c7488d4 _portable_lib.cpython-310-darwin.so`invocation function for block in (anonymous namespace)::copy(MLMultiArray*, executorchcoreml::MultiArray&) + 376 frame pytorch#13: 0x000000030c748ac8 _portable_lib.cpython-310-darwin.so`invocation function for block in (anonymous namespace)::copy(MLMultiArray*, executorchcoreml::MultiArray&) + 52 frame pytorch#14: 0x000000019ad33f4c CoreML`CoreML::MultiArrayBuffer::getBytesWithHandler(void (void const*, unsigned long) block_pointer) const + 340 frame pytorch#15: 0x000000019ad34138 CoreML`-[MLMultiArray(ScopedBufferAccess) getBytesWithHandler:] + 152 frame pytorch#16: 0x000000030c7485ec _portable_lib.cpython-310-darwin.so`(anonymous namespace)::copy(MLMultiArray*, executorchcoreml::MultiArray&) + 296 frame pytorch#17: 0x000000030c744f68 _portable_lib.cpython-310-darwin.so`(anonymous namespace)::set_outputs(std::__1::vector<executorchcoreml::MultiArray, std::__1::allocator<executorchcoreml::MultiArray>>&, NSArray<MLMultiArray*>*) + 180 ``` With this PR, the process succeeds.
quic-boyuc
added a commit
that referenced
this pull request
May 18, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc
added a commit
that referenced
this pull request
Jul 1, 2026
… region
Adds the third top-level Region in the typical AOT+device flow:
Session "<script-name>"
├── quantization/ (pipeline_graph_collector, AOT)
├── edge/ (pipeline_graph_collector, AOT)
└── device/ (this lens; on-device runtime work)
├── adb.execute #1
├── adb.execute #2
└── ...
Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.
Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.
Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.
tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
call triggers it, multiple records share the one region,
disabled-lens does not open it, on_session_end closes the
ExitStack, idempotent _ensure_device_region.
Verification:
source ~/executorch/.venv-11/bin/activate
source ~/executorch/qaisw-*/bin/envsetup.sh
export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
PYTHONPATH=~ python -m pytest \
~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
-> 19 passed (6 new + 13 existing AdbLens tests).
Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc
pushed a commit
that referenced
this pull request
Aug 25, 2026
…2095) ### Summary ETDump records an intermediate tensor by handing the tensor's data pointer to a data sink, and every sink reads those bytes with a plain host read: ```cpp memcpy(cur_data_begin, ptr, length); ``` When the tensor lives on an accelerator that pointer is not host memory, so the read segfaults. A program placed on CUDA crashes as soon as tracing is turned on, which is exactly when someone is trying to debug it. Returning an error instead would not help. All four callers wrap the result in `ET_CHECK_MSG`, so an error aborts the process rather than skipping the tensor. This change brings the data back to host memory first. When the tensor is not on CPU, ETDump looks up the allocator registered for that device type, stages the bytes into a temporary host buffer with `copy_device_to_host`, writes that buffer to the sink and frees it. A tensor on CPU keeps the old path and copies nothing extra. If no allocator is registered for the device, ETDump now reports `NotFound` and logs the device type instead of reading the pointer anyway. ### Test plan Two new test files, each with a CMake target and a Buck target. `devtools/etdump/tests/etdump_device_test.cpp` registers the existing `MockCudaAllocator`, which backs its device memory with host memory, and has two tests. One logs a tensor tagged as CUDA and checks both that ETDump went through the allocator and that the bytes reached the debug buffer. The other logs a tensor on CPU and checks that the allocator was not used at all, so the CPU path is unchanged. `devtools/etdump/tests/etdump_device_no_allocator_test.cpp` covers the case where nothing is registered for the device. The registry is a process wide static with no way to remove an entry, so that case needs a binary that never registers anything, which is why it is a second file. `devtools/etdump/tests/CMakeLists.txt` was not referenced by any parent `CMakeLists.txt`, so nothing in that directory was built by CMake. This adds `add_subdirectory(tests)` to `devtools/etdump/CMakeLists.txt` under `BUILD_TESTING`, so both new tests are picked up by `ctest`, which is how the C++ tests run. The pre-existing `sdk_etdump_tests` target stays out of the CMake build. It compiles `etdump_test.cpp`, which includes `etdump_filter.h`, which needs re2, and the devtools build does not pull re2 in. It is now guarded on re2 being available rather than being silently unreachable. With this change both new binaries pass under `ctest`: ``` 1/2 Test #1: etdump_device_test ................ Passed 2/2 Test #2: etdump_device_no_allocator_test ... Passed ``` With `etdump_flatcc.cpp` reverted to the old code and everything rebuilt, both fail: ``` Expected equality of these values: g_mock_cuda.d2h_count_ Which is: 0 1 Death test: etdump_gen.log_evalue(EValue(tensor)) Result: failed to die. ``` Also reproduced the real crash on one NVIDIA H100, with a small program that allocates through `cudaMalloc`, tags a tensor as CUDA and logs it: ``` before: Segmentation fault (core dumped) after: debug buffer holds 1.5 2.5 3.5 4.5 ``` Checked that `etdump_flatcc.cpp` still compiles with `-DUSE_ATEN_LIB`. `clang-format` reports no changes needed on the four touched C++ files. ### Landing order pytorch#22058 adds device-planned arenas to the Python bindings that ETDump can be handed pointers into, and without this change `BufferDataSink::write` would `memcpy` device memory from the host. This should land before or together with it. ### Not covered The ATen mode branch of the device type conversion only compiles. There is no ATen mode CMake build to run it in, so nothing here executes it. It is reachable in principle: `runtime/executor/tensor_parser_aten.cpp` reads the serialized device type and index, including CUDA, and builds the ATen tensor on that device. An earlier version of this description said ATen tensors carry no device metadata, which is wrong. `LogTensorOnCpuDoesNotStageThroughTheAllocator` passes with the production change reverted, since CPU tensors already went straight to the data sink. It documents the CPU path rather than locking the fix; the other two tests are the ones that require the new branch. --------- Co-authored-by: r <r@e>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Please DO NOT MERGE this PR.
This PR is to reproduce the issue where nn.Module inference is working fine, however, once after
torch.export.export, sqnr went from 120 -> 8.Command:
python examples/qualcomm/oss_scripts/moshi/mimi.py -b build-android -s $device -m SM8650 --chunks_per_batch 125To run nn.Module where it is working fine, please uncomment line 264-267 and comment out line 269-273