Skip to content

Dev1/winskuo/gh /aihub remove code - #2

Closed
winskuo-quic wants to merge 21 commits into
mainfrom
dev1/winskuo/gh_/aihub_remove_code
Closed

Dev1/winskuo/gh /aihub remove code#2
winskuo-quic wants to merge 21 commits into
mainfrom
dev1/winskuo/gh_/aihub_remove_code

Conversation

@winskuo-quic

Copy link
Copy Markdown
Collaborator

Summary

[PLEASE REMOVE] See CONTRIBUTING.md's Pull Requests for ExecuTorch PR guidelines.

[PLEASE REMOVE] If this PR closes an issue, please add a Fixes #<issue-id> line.

[PLEASE REMOVE] If this PR introduces a fix or feature that should be the upcoming release notes, please add a "Release notes: " label. For a list of available release notes labels, check out CONTRIBUTING.md's Pull Requests.

Test plan

[PLEASE REMOVE] How did you test this PR? Please write down any manual commands you used and note down tests that you have written if applicable.

JacobSzwejbka and others added 21 commits April 24, 2026 15:25
Summary:

Attempt 3 to check numel and nbytes overflow. This time we defer
checking dynamic sized inputs until their size is realized.

Reviewed By: lucylq

Differential Revision: D98148157
…inear (pytorch#19117)

## Summary

The bilinear grid_sampler_2d portable kernel computes interpolation
weights via subtractions like `(ix_se - ix)` where both operands are
close integer-valued coordinates in pixel space. In fp16 (10 bits of
mantissa) that's classic catastrophic cancellation — the result has only
a handful of significant bits. The downstream weighted-sum accumulation
then loses further precision.

Measured on a unit test exercising interior grid points with fp16
inputs, the kernel drifts by ~0.1 absolute from an fp32 reference.
That's visible as incorrect depth / flow output near non-integer sample
points, which is most of them.

## Fix

An `AccType<CTYPE>` trait mapping `Half` and `BFloat16` to `float`,
leaving every other dtype unchanged. Used for intermediate coordinate,
weight computation, and `out_val` accumulation. Loads cast `CTYPE ->
ACC`; the final store casts `ACC -> CTYPE` once. Only internal math is
promoted; memory layout / public API / tensor dtypes are unchanged.

```cpp
template <typename CTYPE>
using AccType = std::conditional_t<
    std::is_same_v<CTYPE, executorch::aten::Half> ||
        std::is_same_v<CTYPE, executorch::aten::BFloat16>,
    float,
    CTYPE>;
```

## Effects

- **fp32 / Int / any non-half dtype**: `AccType<T>` is `T`, so the
generated code is byte-identical. No behavior change.
- **Half / BFloat16**: `max_abs` vs an fp32 reference drops from **~0.1
to 0** on the shapes I tested (N=1..2, C=7..64, H/W up to 96, both
`align_corners` values).
- **Perf**: a handful of fp16↔fp32 conversions per output element. Not
measurable at op level; well within the portable kernel's scalar cost
envelope.

## Scope

Only touches the bilinear interpolation path. The nearest-mode path
doesn't do weighted-sum accumulation and doesn't have the cancellation
issue — left alone in this change.

## Test plan

- [x] Builds clean for Android arm64 and host (Apple Clang 21).
- [x] Verified numerically via a standalone harness that runs the kernel
with matched fp32 / fp16 inputs and compares against an
fp32-then-downcast reference. All shapes pass within a single fp16 ULP
(or are bit-exact). fp32 tests remain bit-identical to the pre-change
kernel.
- [x] Existing `kernels/test/op_grid_sampler_2d_test.cpp` unit tests
continue to pass (both fp32 shapes that were previously tested, and the
fp16 path I'm specifically fixing).

Happy to add an fp16-specific test case to `op_grid_sampler_2d_test.cpp`
if useful for CI coverage here — just let me know the preferred
approach.

cc @larryliu0820 @manuelcandales
Differential Revision: D101728720

Pull Request resolved: pytorch#19052
### Summary
Fixes pytorch#18924
Extends `aten.bitwise_not` support in the MLX delegate to handle integer
tensors, not just boolean tensors.
Previously the handler only dispatched to `LogicalNotNode` for `bool`
and raised `NotImplementedError` for all other dtypes. This adds a
dedicated `BitwiseInvertNode` backed by `mlx::core::bitwise_invert`, and
updates the handler to dispatch based on dtype:
 - `bool` → `LogicalNotNode` (unchanged)
 - `int32`, `int64` → `BitwiseInvertNode`
#### Changes:
- `serialization/schema.fbs`: add `BitwiseInvertNode` table and append
to `OpNode` union
- `runtime/MLXInterpreter.h`: add `exec_bitwise_invert()` and dispatch
case
- `ops.py`: update `_bitwise_not_handler` to dispatch to
`BitwiseInvertNode` for integers
- `test/test_ops.py`: add `bitwise_not_int` test for `int32` and `int64`
### Test plan
All tests were ran on a machine with an Apple M1 Pro CPU, macOS 26.4.1.
- `python3 -m py_compile backends/mlx/ops.py
backends/mlx/test/test_ops.py`
- `python3 backends/mlx/serialization/generate.py`
- `python3 -m executorch.backends.mlx.test.run_all_tests
bitwise_not_int`
### Test output
```
============================================================
TEST SUMMARY
============================================================
Passed: 6
Failed: 0
============================================================
```
cc @metascroy

---------

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Scott Roy <161522778+metascroy@users.noreply.github.com>
Differential Revision: D102070375

Pull Request resolved: pytorch#19074
Add Closeable interface so Module can be used with try-with-resources.
close() delegates to destroy(). Also make destroy() idempotent by
checking mHybridData.isValid() before calling resetNative(), satisfying
the Closeable contract.

This commit was authored with the help of Claude.
TrainingModule: implement Closeable, replace Log.e + silent empty
returns with IllegalStateException throws. Add checkNotDestroyed() guard
on all public methods.

SGD: throw IllegalStateException instead of bare RuntimeException when
optimizer is destroyed.

AsrModule: throw ExecutorchRuntimeException instead of bare
RuntimeException on transcription failure.

ExecuTorchRuntime.validateFilePath: throw IllegalArgumentException
instead of bare RuntimeException, with descriptive message.

JNI constructors: wrap ExecuTorchJni and ExecuTorchLlmJni constructor
bodies in try-catch so C++ exceptions become ExecutorchRuntimeException
instead of generic RuntimeException.

This commit was authored with the help of Claude.

---------

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Differential Revision: D102425537

Pull Request resolved: pytorch#19128
…_nhwc operator + tests

Differential Revision: D96507563

Pull Request resolved: pytorch#18479
…ch#19028 (pytorch#19133)

## Summary

Reverts the following Android PRs:
- pytorch#19099 — Android: consistent error types across all modules
- pytorch#19124 — Android: Module implements Closeable
- pytorch#19092 — Android: improve error diagnostics for LlmModule and
exceptions
- pytorch#19028 — Ignored Module tests: provide required input tensor

Authored with Claude.
Differential Revision: D102488314

Pull Request resolved: pytorch#19134
Differential Revision: D102385104

Pull Request resolved: pytorch#19122
Differential Revision: D102493794

Pull Request resolved: pytorch#19136
@winskuo-quic
winskuo-quic changed the base branch from dev1/winskuo/gh_aihub_remove_readme to main April 27, 2026 02:47
@winskuo-quic
winskuo-quic deleted the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:47
@winskuo-quic
winskuo-quic restored the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:48
@winskuo-quic
winskuo-quic deleted the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:49
@winskuo-quic
winskuo-quic restored the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:49
@winskuo-quic
winskuo-quic deleted the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:50
@winskuo-quic
winskuo-quic restored the dev1/winskuo/gh_/aihub_remove_code branch April 27, 2026 02:50
quic-boyuc added a commit that referenced this pull request May 5, 2026
Add observe_pass, a decorator that wraps any PassBase subclass or callable
pass to automatically collect input/output graphs via Observatory. Names
are derived from the class/function name and deduplicated by collect()
itself (#2, pytorch#3, ...) so repeated calls never silently overwrite records.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request May 18, 2026
… region

Adds the third top-level Region in the typical AOT+device flow:

    Session "<script-name>"
    ├── quantization/   (pipeline_graph_collector, AOT)
    ├── edge/           (pipeline_graph_collector, AOT)
    └── device/         (this lens; on-device runtime work)
          ├── adb.execute #1
          ├── adb.execute #2
          └── ...

Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.

Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.

Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.

tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
  call triggers it, multiple records share the one region,
  disabled-lens does not open it, on_session_end closes the
  ExitStack, idempotent _ensure_device_region.

Verification:
  source ~/executorch/.venv-11/bin/activate
  source ~/executorch/qaisw-*/bin/envsetup.sh
  export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
  PYTHONPATH=~ python -m pytest \
    ~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
  -> 19 passed (6 new + 13 existing AdbLens tests).

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request May 22, 2026
In compare mode each archive column already shows the label in its sticky
header (commit 5ae2a72), so the `<label>/` that compare_archives prepends
to every record and session name renders as redundant noise inside the
column and squeezes the natural-width layout from c9e7bda.

Strip is display-only — `rec.name` / `session.name` in `state.data` stay
prefixed so graph_ref resolution and the global record map keep working.
A `displayPrefix` arg threads through `renderRecordItem`,
`_renderSessionDashboardLink`, `renderIndexFlat`, and `renderIndexTree`;
`_renderArchiveColumn` passes `archive.label`. Single-archive paths pass
no prefix and are unchanged. Collision suffixes (`A/foo #2`) strip cleanly
to `foo #2` since only the head matches.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
Add observe_pass, a decorator that wraps any PassBase subclass or callable
pass to automatically collect input/output graphs via Observatory. Names
are derived from the class/function name and deduplicated by collect()
itself (#2, pytorch#3, ...) so repeated calls never silently overwrite records.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
… region

Adds the third top-level Region in the typical AOT+device flow:

    Session "<script-name>"
    ├── quantization/   (pipeline_graph_collector, AOT)
    ├── edge/           (pipeline_graph_collector, AOT)
    └── device/         (this lens; on-device runtime work)
          ├── adb.execute #1
          ├── adb.execute #2
          └── ...

Region opens lazily on the first patched SimpleADB operation
(via `_ensure_device_region`, called from `note_simple_adb`), so a
CLI run with no on-device inference doesn't get an empty group
cluttering the tree view. The region is closed in `on_session_end`
via a `contextlib.ExitStack`.

Each inference record (`adb.execute #N`) lives **directly** under
`device` — no per-call sub-region — so the region holds multiple
records as designed in RFC §4.5.

Disabled lens (`config={"adb": {"enabled": False}}`) skips region
opening entirely.

tests/test_adb_lens_regions.py (new):
- 6 tests covering: lazy device-region opening, first SimpleADB
  call triggers it, multiple records share the one region,
  disabled-lens does not open it, on_session_end closes the
  ExitStack, idempotent _ensure_device_region.

Verification:
  source ~/executorch/.venv-11/bin/activate
  source ~/executorch/qaisw-*/bin/envsetup.sh
  export QNN_SDK_ROOT=~/executorch/qaisw-v2.41.0.251111144900_191416
  PYTHONPATH=~ python -m pytest \
    ~/executorch/backends/qualcomm/debugger/observatory/tests/ -v
  -> 19 passed (6 new + 13 existing AdbLens tests).

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
quic-boyuc added a commit that referenced this pull request Jul 1, 2026
In compare mode each archive column already shows the label in its sticky
header (commit 5ae2a72), so the `<label>/` that compare_archives prepends
to every record and session name renders as redundant noise inside the
column and squeezes the natural-width layout from c9e7bda.

Strip is display-only — `rec.name` / `session.name` in `state.data` stay
prefixed so graph_ref resolution and the global record map keep working.
A `displayPrefix` arg threads through `renderRecordItem`,
`_renderSessionDashboardLink`, `renderIndexFlat`, and `renderIndexTree`;
`_renderArchiveColumn` passes `archive.label`. Single-archive paths pass
no prefix and are unchanged. Collision suffixes (`A/foo #2`) strip cleanly
to `foo #2` since only the head matches.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
qti-horodnic pushed a commit that referenced this pull request Jul 28, 2026
### Summary
- fill the test gap for newly added op
- utils test

### Test plan
pytest backends/qualcomm/tests/rework/utils/test.py
quic-boyuc pushed a commit that referenced this pull request Aug 25, 2026
…2095)

### Summary

ETDump records an intermediate tensor by handing the tensor's data
pointer to a
data sink, and every sink reads those bytes with a plain host read:

```cpp
memcpy(cur_data_begin, ptr, length);
```

When the tensor lives on an accelerator that pointer is not host memory,
so the
read segfaults. A program placed on CUDA crashes as soon as tracing is
turned
on, which is exactly when someone is trying to debug it.

Returning an error instead would not help. All four callers wrap the
result in
`ET_CHECK_MSG`, so an error aborts the process rather than skipping the
tensor.

This change brings the data back to host memory first. When the tensor
is not on
CPU, ETDump looks up the allocator registered for that device type,
stages the
bytes into a temporary host buffer with `copy_device_to_host`, writes
that buffer
to the sink and frees it. A tensor on CPU keeps the old path and copies
nothing
extra.

If no allocator is registered for the device, ETDump now reports
`NotFound` and
logs the device type instead of reading the pointer anyway.

### Test plan

Two new test files, each with a CMake target and a Buck target.

`devtools/etdump/tests/etdump_device_test.cpp` registers the existing
`MockCudaAllocator`, which backs its device memory with host memory, and
has two
tests. One logs a tensor tagged as CUDA and checks both that ETDump went
through
the allocator and that the bytes reached the debug buffer. The other
logs a
tensor on CPU and checks that the allocator was not used at all, so the
CPU path
is unchanged.

`devtools/etdump/tests/etdump_device_no_allocator_test.cpp` covers the
case where
nothing is registered for the device. The registry is a process wide
static with
no way to remove an entry, so that case needs a binary that never
registers
anything, which is why it is a second file.

`devtools/etdump/tests/CMakeLists.txt` was not referenced by any parent
`CMakeLists.txt`, so nothing in that directory was built by CMake. This
adds
`add_subdirectory(tests)` to `devtools/etdump/CMakeLists.txt` under
`BUILD_TESTING`, so both new tests are picked up by `ctest`, which is
how the C++
tests run.

The pre-existing `sdk_etdump_tests` target stays out of the CMake build.
It
compiles `etdump_test.cpp`, which includes `etdump_filter.h`, which
needs re2,
and the devtools build does not pull re2 in. It is now guarded on re2
being
available rather than being silently unreachable.

With this change both new binaries pass under `ctest`:

```
1/2 Test #1: etdump_device_test ................   Passed
2/2 Test #2: etdump_device_no_allocator_test ...   Passed
```

With `etdump_flatcc.cpp` reverted to the old code and everything
rebuilt, both
fail:

```
Expected equality of these values:
  g_mock_cuda.d2h_count_
    Which is: 0
  1

Death test: etdump_gen.log_evalue(EValue(tensor))
    Result: failed to die.
```

Also reproduced the real crash on one NVIDIA H100, with a small program
that
allocates through `cudaMalloc`, tags a tensor as CUDA and logs it:

```
before: Segmentation fault (core dumped)
after:  debug buffer holds 1.5 2.5 3.5 4.5
```

Checked that `etdump_flatcc.cpp` still compiles with `-DUSE_ATEN_LIB`.

`clang-format` reports no changes needed on the four touched C++ files.

### Landing order

pytorch#22058 adds device-planned arenas to the Python bindings that ETDump can
be handed
pointers into, and without this change `BufferDataSink::write` would
`memcpy` device
memory from the host. This should land before or together with it.

### Not covered

The ATen mode branch of the device type conversion only compiles. There
is no
ATen mode CMake build to run it in, so nothing here executes it. It is
reachable
in principle: `runtime/executor/tensor_parser_aten.cpp` reads the
serialized
device type and index, including CUDA, and builds the ATen tensor on
that device.
An earlier version of this description said ATen tensors carry no device
metadata,
which is wrong.

`LogTensorOnCpuDoesNotStageThroughTheAllocator` passes with the
production change
reverted, since CPU tensors already went straight to the data sink. It
documents
the CPU path rather than locking the fix; the other two tests are the
ones that
require the new branch.

---------

Co-authored-by: r <r@e>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

10 participants