Skip to content

[DO NOT MERGE]Qualcomm AI Engine Direct - Mimi Decoder low SQNR reproduce - #1

Draft
winskuo-quic wants to merge 1 commit into
dev1/winskuo/mimi_stage2from
dev1/winskuo/mimi_streaming_low_sqnr_reproduce
Draft

[DO NOT MERGE]Qualcomm AI Engine Direct - Mimi Decoder low SQNR reproduce#1
winskuo-quic wants to merge 1 commit into
dev1/winskuo/mimi_stage2from
dev1/winskuo/mimi_streaming_low_sqnr_reproduce

Conversation

@winskuo-quic

Copy link
Copy Markdown
Collaborator

Summary

Please DO NOT MERGE this PR.
This PR is to reproduce the issue where nn.Module inference is working fine, however, once after torch.export.export, sqnr went from 120 -> 8.

Command:
python examples/qualcomm/oss_scripts/moshi/mimi.py -b build-android -s $device -m SM8650 --chunks_per_batch 125
To run nn.Module where it is working fine, please uncomment line 264-267 and comment out line 269-273

@winskuo-quic
winskuo-quic changed the base branch from dev1/winskuo/mimi_stage2 to main April 15, 2025 03:18
@winskuo-quic
winskuo-quic changed the base branch from main to dev1/winskuo/mimi_stage2 April 15, 2025 03:19
chenweng-quic pushed a commit that referenced this pull request Jun 9, 2025
Differential Revision: D75911655

Pull Request resolved: pytorch#11344
shewu-quic pushed a commit that referenced this pull request Aug 1, 2025
BNNS copy crashes the process when the dtypes differ
(pytorch#11714).

With the example in this PR
(pytorch#11714), we crash the
process on main. Here is the stack trace from LLDB:

```
Process 19234 stopped
* thread #1, queue = 'com.apple.main-thread', stop reason = signal SIGABRT
    frame #0: 0x0000000190ac9388 libsystem_kernel.dylib`__pthread_kill + 8
libsystem_kernel.dylib`__pthread_kill:
->  0x190ac9388 <+8>:  b.lo   0x190ac93a8    ; <+40>
    0x190ac938c <+12>: pacibsp 
    0x190ac9390 <+16>: stp    x29, x30, [sp, #-0x10]!
    0x190ac9394 <+20>: mov    x29, sp
(lldb) bt
* thread #1, queue = 'com.apple.main-thread', stop reason = signal SIGABRT
  * frame #0: 0x0000000190ac9388 libsystem_kernel.dylib`__pthread_kill + 8
    frame #1: 0x0000000190b0288c libsystem_pthread.dylib`pthread_kill + 296
    frame #2: 0x0000000190a0bc60 libsystem_c.dylib`abort + 124
    frame pytorch#3: 0x0000000190910174 libsystem_malloc.dylib`malloc_vreport + 892
    frame pytorch#4: 0x0000000190913c90 libsystem_malloc.dylib`malloc_report + 64
    frame pytorch#5: 0x000000019091821c libsystem_malloc.dylib`___BUG_IN_CLIENT_OF_LIBMALLOC_POINTER_BEING_FREED_WAS_NOT_ALLOCATED + 32
    frame pytorch#6: 0x000000019d2f4084 libBNNS.dylib`___lldb_unnamed_symbol1620 + 564
    frame pytorch#7: 0x000000019d2f5bac libBNNS.dylib`___lldb_unnamed_symbol1628 + 680
    frame pytorch#8: 0x000000019d69ce48 libBNNS.dylib`BNNSCopy + 616
    frame pytorch#9: 0x000000030c74d950 _portable_lib.cpython-310-darwin.so`(anonymous namespace)::copy_using_bnns(executorchcoreml::MultiArray const&, executorchcoreml::MultiArray&) + 188
    frame pytorch#10: 0x000000030c74cfdc _portable_lib.cpython-310-darwin.so`(anonymous namespace)::copy(executorchcoreml::MultiArray const&, executorchcoreml::MultiArray&, executorchcoreml::MultiArray::CopyOptions) + 72
    frame pytorch#11: 0x000000030c74ceec _portable_lib.cpython-310-darwin.so`executorchcoreml::MultiArray::copy(executorchcoreml::MultiArray&, executorchcoreml::MultiArray::CopyOptions) const + 148
    frame pytorch#12: 0x000000030c7488d4 _portable_lib.cpython-310-darwin.so`invocation function for block in (anonymous namespace)::copy(MLMultiArray*, executorchcoreml::MultiArray&) + 376
    frame pytorch#13: 0x000000030c748ac8 _portable_lib.cpython-310-darwin.so`invocation function for block in (anonymous namespace)::copy(MLMultiArray*, executorchcoreml::MultiArray&) + 52
    frame pytorch#14: 0x000000019ad33f4c CoreML`CoreML::MultiArrayBuffer::getBytesWithHandler(void (void const*, unsigned long) block_pointer) const + 340
    frame pytorch#15: 0x000000019ad34138 CoreML`-[MLMultiArray(ScopedBufferAccess) getBytesWithHandler:] + 152
    frame pytorch#16: 0x000000030c7485ec _portable_lib.cpython-310-darwin.so`(anonymous namespace)::copy(MLMultiArray*, executorchcoreml::MultiArray&) + 296
    frame pytorch#17: 0x000000030c744f68 _portable_lib.cpython-310-darwin.so`(anonymous namespace)::set_outputs(std::__1::vector<executorchcoreml::MultiArray, std::__1::allocator<executorchcoreml::MultiArray>>&, NSArray<MLMultiArray*>*) + 180
```


With this PR, the process succeeds.
quic-boyuc added a commit that referenced this pull request May 18, 2026
… region

Adds the third top-level Region in the typical AOT+device flow:

    Session "<script-name>"
    ├── quantization/   (pipeline_graph_collector, AOT)
    ├── edge/           (pipeline_graph_collector, AOT)
    └── device/         (this lens; on-device runtime work)
          ├── adb.execute #1
          ├── adb.execute #2
          └── ...

Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.

Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.

Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.

tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
  call triggers it, multiple records share the one region,
  disabled-lens does not open it, on_session_end closes the
  ExitStack, idempotent _ensure_device_region.

Verification:
  source ~/executorch/.venv-11/bin/activate
  source ~/executorch/qaisw-*/bin/envsetup.sh
  export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
  PYTHONPATH=~ python -m pytest \
    ~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
  -> 19 passed (6 new + 13 existing AdbLens tests).

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
… region

Adds the third top-level Region in the typical AOT+device flow:

    Session "<script-name>"
    ├── quantization/   (pipeline_graph_collector, AOT)
    ├── edge/           (pipeline_graph_collector, AOT)
    └── device/         (this lens; on-device runtime work)
          ├── adb.execute #1
          ├── adb.execute #2
          └── ...

Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.

Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.

Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.

tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
  call triggers it, multiple records share the one region,
  disabled-lens does not open it, on_session_end closes the
  ExitStack, idempotent _ensure_device_region.

Verification:
  source ~/executorch/.venv-11/bin/activate
  source ~/executorch/qaisw-*/bin/envsetup.sh
  export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
  PYTHONPATH=~ python -m pytest \
    ~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
  -> 19 passed (6 new + 13 existing AdbLens tests).

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc pushed a commit that referenced this pull request Aug 25, 2026
…2095)

### Summary

ETDump records an intermediate tensor by handing the tensor's data
pointer to a
data sink, and every sink reads those bytes with a plain host read:

```cpp
memcpy(cur_data_begin, ptr, length);
```

When the tensor lives on an accelerator that pointer is not host memory,
so the
read segfaults. A program placed on CUDA crashes as soon as tracing is
turned
on, which is exactly when someone is trying to debug it.

Returning an error instead would not help. All four callers wrap the
result in
`ET_CHECK_MSG`, so an error aborts the process rather than skipping the
tensor.

This change brings the data back to host memory first. When the tensor
is not on
CPU, ETDump looks up the allocator registered for that device type,
stages the
bytes into a temporary host buffer with `copy_device_to_host`, writes
that buffer
to the sink and frees it. A tensor on CPU keeps the old path and copies
nothing
extra.

If no allocator is registered for the device, ETDump now reports
`NotFound` and
logs the device type instead of reading the pointer anyway.

### Test plan

Two new test files, each with a CMake target and a Buck target.

`devtools/etdump/tests/etdump_device_test.cpp` registers the existing
`MockCudaAllocator`, which backs its device memory with host memory, and
has two
tests. One logs a tensor tagged as CUDA and checks both that ETDump went
through
the allocator and that the bytes reached the debug buffer. The other
logs a
tensor on CPU and checks that the allocator was not used at all, so the
CPU path
is unchanged.

`devtools/etdump/tests/etdump_device_no_allocator_test.cpp` covers the
case where
nothing is registered for the device. The registry is a process wide
static with
no way to remove an entry, so that case needs a binary that never
registers
anything, which is why it is a second file.

`devtools/etdump/tests/CMakeLists.txt` was not referenced by any parent
`CMakeLists.txt`, so nothing in that directory was built by CMake. This
adds
`add_subdirectory(tests)` to `devtools/etdump/CMakeLists.txt` under
`BUILD_TESTING`, so both new tests are picked up by `ctest`, which is
how the C++
tests run.

The pre-existing `sdk_etdump_tests` target stays out of the CMake build.
It
compiles `etdump_test.cpp`, which includes `etdump_filter.h`, which
needs re2,
and the devtools build does not pull re2 in. It is now guarded on re2
being
available rather than being silently unreachable.

With this change both new binaries pass under `ctest`:

```
1/2 Test #1: etdump_device_test ................   Passed
2/2 Test #2: etdump_device_no_allocator_test ...   Passed
```

With `etdump_flatcc.cpp` reverted to the old code and everything
rebuilt, both
fail:

```
Expected equality of these values:
  g_mock_cuda.d2h_count_
    Which is: 0
  1

Death test: etdump_gen.log_evalue(EValue(tensor))
    Result: failed to die.
```

Also reproduced the real crash on one NVIDIA H100, with a small program
that
allocates through `cudaMalloc`, tags a tensor as CUDA and logs it:

```
before: Segmentation fault (core dumped)
after:  debug buffer holds 1.5 2.5 3.5 4.5
```

Checked that `etdump_flatcc.cpp` still compiles with `-DUSE_ATEN_LIB`.

`clang-format` reports no changes needed on the four touched C++ files.

### Landing order

pytorch#22058 adds device-planned arenas to the Python bindings that ETDump can
be handed
pointers into, and without this change `BufferDataSink::write` would
`memcpy` device
memory from the host. This should land before or together with it.

### Not covered

The ATen mode branch of the device type conversion only compiles. There
is no
ATen mode CMake build to run it in, so nothing here executes it. It is
reachable
in principle: `runtime/executor/tensor_parser_aten.cpp` reads the
serialized
device type and index, including CUDA, and builds the ATen tensor on
that device.
An earlier version of this description said ATen tensors carry no device
metadata,
which is wrong.

`LogTensorOnCpuDoesNotStageThroughTheAllocator` passes with the
production change
reverted, since CPU tensors already went straight to the data sink. It
documents
the CPU path rather than locking the fix; the other two tests are the
ones that
require the new branch.

---------

Co-authored-by: r <r@e>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant