Qualcomm AI Engine Direct - [GenAI Pipeline] PR4: Adapter interfaces, default implementations & dataset providers - #21751
Conversation
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21751
Note: Links to docs will display an error until the docs builds have been completed. ✅ No FailuresAs of commit 3aac96f with merge base c56e6bf ( This comment was automatically generated by Dr. CI and updates every 15 minutes. |
|
@pytorchbot label "release notes: qualcomm" |
|
@qti-horodnic Firstly, this is a great refactor, modular, composable, and creating abstractions for all external |
|
|
@psiddh Phase 1: The current work, includes PRs 1-7 as outlined in the PR description. Includes only the addition of the skeleton of the new GenAI infrastructure. No existing code paths change behavior, and the skeleton isn't wired into any current entry point. This phase is purely additive and inert by construction, which is why several interfaces here are single-graph / stub-bodied: they're the N=1 degenerate case, with the general form landing in Phase 2. Phase 2: The next phase will include 4 PRs, divided between me and @DannyYuyang-quic into 2 (roughly) parallel work streams. This phase will include moving legacy code (e.g. Phase 3: Cleanup of old code. This will be done after only phase 2 has been completed and all critical code paths have been validated successfully. Nothing is deleted before this phase. |
On the HF -> static flat-KV swap: I'd keep that out of the transform list since it's "which class to construct" and is already registry data. Separating the two keeps the
|
…and default implementations
|
Merging it now (inert for now) , as it unblocks the next few PRs |
… quantization strategy implementations (#21899) ## Summary This PR implements the __model preparation__ and __quantization__ strategy implementations, replacing the `NotImplementedError` stubs with real logic. Each strategy delegates to injectable adapter interfaces (from PR4) for testability. ### What's included #### Strategy implementations (2 files + 1 `__init__` fix): - `ExecuTorchModelPreparationStrategy`: 5-step flow - `load_model` → `load_tokenizer` (via `ModelLoaderAdapter`) - `generate_calibration_data` (via separately-injectable `CalibrationDataAdapter`) - Optional tokenizer export for on-device runtime - Chat template extraction from tokenizer (with `extra_options` fallback) - Validates input config (`model_name`, `soc_model` required) - `ExecuTorchQuantizationStrategy`: Full PT2E single-graph pipeline via `QuantizerAdapter` - export → make_quantizer → prepare_pt2e → calibrate → convert_pt2e - Supports `quant_dtype`, `quant_recipe`, and per-channel options via `extra_options` - Handles any `Iterable` as calibration data (lists, DataLoaders, generators) - Validates calibration data is non-empty before export - Warns (does not fail) when `training_data` is provided (QAT deferred) - `strategies/model_preparation/__init__.py`: adds missing `ExecuTorchModelPreparationStrategy` import to `__all__` #### Unit tests: - `test_executorch_model_preparation_strategy.py` - `test_executorch_quantization_strategy.py` - `test_default_model_preparation_adapter.py` - `test_default_model_preparation_adapter.py` ### PR Review Checklist - All new classes follow single responsibility (one class per file) - Yes. - All dependencies are injected via constructor with sensible defaults - Yes. - All external calls are behind injectable interfaces - Yes (adapter pattern). - Unit tests cover every public method - Yes (100% coverage on strategy impls). - Docstrings on all public classes and methods - Yes. - Type annotations on all function signatures - Yes. - Logging follows the strategy in the LLD - Yes (info on entry/exit, debug per step). ### Related PRs - PR 1: Core data model, engine routing & exceptions: #20409 - PR 2: Strategy interfaces & stage wrappers: #20795 - PR 3: Pipeline orchestrator: #21149 - PR 4: Adapter interfaces, default implementations & dataset providers: #21751 - PR 5: Model preparation & quantization strategies: this pr. - PR 6: Compilation & inference strategy implementations: pending. - PR 7: Integration & E2E tests: pending. ## Test plan ### Run only tests added in this PR: ``` python -m pytest \ backends/qualcomm/genai_pipeline/tests/strategies/model_preparation/ \ backends/qualcomm/genai_pipeline/tests/strategies/quantization/ \ -v ``` ### Run only this PR's tests with coverage: ``` python -m pytest \ backends/qualcomm/genai_pipeline/tests/strategies/model_preparation/ \ backends/qualcomm/genai_pipeline/tests/strategies/quantization/ \ --cov=backends/qualcomm/genai_pipeline/strategies/model_preparation \ --cov=backends/qualcomm/genai_pipeline/strategies/quantization \ --cov-config=backends/qualcomm/.coveragerc \ --cov-report=term-missing ``` Result: ``` Name Stmts Miss Branch BrPart Cover Missing ---------------------------------------------------------------------------------------------------------------------------------------------------- backends/qualcomm/genai_pipeline/strategies/model_preparation/executorch_model_preparation_strategy.py 68 0 14 0 100% backends/qualcomm/genai_pipeline/strategies/model_preparation/model_loader_adapter.py 9 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/model_preparation/model_preparation_strategy.py 7 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/quantization/executorch_quantization_strategy.py 61 0 18 0 100% backends/qualcomm/genai_pipeline/strategies/quantization/quantization_strategy.py 7 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/quantization/quantizer_adapter.py 9 0 0 0 100% ---------------------------------------------------------------------------------------------------------------------------------------------------- TOTAL 161 0 32 0 100% ``` ### Run all `genai_pipeline` tests: ``` python -m pytest backends/qualcomm/genai_pipeline/tests/ -v ``` ### Run all `genai_pipeline` tests with coverage: ``` python -m pytest backends/qualcomm/genai_pipeline/tests/ \ --cov=backends/qualcomm/genai_pipeline \ --cov-config=backends/qualcomm/.coveragerc \ --cov-report=term-missing ``` Result: ``` Name Stmts Miss Branch BrPart Cover Missing ---------------------------------------------------------------------------------------------------------------------------------------------------- backends/qualcomm/genai_pipeline/configs/compilation_input_config.py 11 0 0 0 100% backends/qualcomm/genai_pipeline/configs/compilation_output_config.py 8 0 0 0 100% backends/qualcomm/genai_pipeline/configs/inference_input_config.py 12 0 0 0 100% backends/qualcomm/genai_pipeline/configs/inference_output_config.py 9 0 0 0 100% backends/qualcomm/genai_pipeline/configs/model_preparation_input_config.py 7 0 0 0 100% backends/qualcomm/genai_pipeline/configs/model_preparation_output_config.py 12 0 0 0 100% backends/qualcomm/genai_pipeline/configs/quantization_input_config.py 13 0 0 0 100% backends/qualcomm/genai_pipeline/configs/quantization_output_config.py 5 0 0 0 100% backends/qualcomm/genai_pipeline/datasets/calibration_data_adapter.py 5 0 0 0 100% backends/qualcomm/genai_pipeline/datasets/default_calibration_data_adapter.py 26 0 6 0 100% backends/qualcomm/genai_pipeline/datasets/default_training_data_adapter.py 14 0 2 0 100% backends/qualcomm/genai_pipeline/datasets/training_data_adapter.py 5 0 0 0 100% backends/qualcomm/genai_pipeline/engine_proxy.py 20 0 4 0 100% backends/qualcomm/genai_pipeline/exceptions.py 20 0 6 0 100% backends/qualcomm/genai_pipeline/genai_pipeline.py 99 7 12 1 93% 190-201 backends/qualcomm/genai_pipeline/pipeline_context.py 52 0 14 0 100% backends/qualcomm/genai_pipeline/pipeline_stage.py 5 0 0 0 100% backends/qualcomm/genai_pipeline/stages/compilation_stage.py 14 0 0 0 100% backends/qualcomm/genai_pipeline/stages/inference_stage.py 14 0 0 0 100% backends/qualcomm/genai_pipeline/stages/model_preparation_stage.py 14 2 0 0 86% 30, 37 backends/qualcomm/genai_pipeline/stages/quantization_stage.py 14 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/compilation/compilation_strategy.py 7 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/compilation/compiler_adapter.py 11 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/compilation/executorch_compilation_strategy.py 7 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/inference/device_runner_adapter.py 14 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/inference/executorch_inference_strategy.py 7 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/inference/inference_strategy.py 7 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/model_preparation/executorch_model_preparation_strategy.py 68 0 14 0 100% backends/qualcomm/genai_pipeline/strategies/model_preparation/model_loader_adapter.py 9 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/model_preparation/model_preparation_strategy.py 7 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/quantization/executorch_quantization_strategy.py 61 0 18 0 100% backends/qualcomm/genai_pipeline/strategies/quantization/quantization_strategy.py 7 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/quantization/quantizer_adapter.py 9 0 0 0 100% ---------------------------------------------------------------------------------------------------------------------------------------------------- TOTAL 593 9 76 1 99% ```
…ence strategy implementations (#22284) ## Summary This PR implements the __compilation__ and __inference__ strategy implementations, completing the strategy layer. Each strategy delegates to injectable adapter interfaces for testability. ### What's included #### Strategy implementations: - `ExecuTorchCompilationStrategy`: Compiles model to .pte artifacts via `CompilerAdapter` - Validates model, example_inputs, soc_model, and backend_type are present - Passes `example_inputs` explicitly to the adapter (not via `extra_options`) — mirrors the PR5 fix for the quantization stage - Delegates to adapter with example_inputs, compile specs, artifact dir, soc_model, backend_type - Filters `context.extra_options` to a compilation-relevant allow-list - Returns artifact paths and optional `ETRecord` - `ExecuTorchInferenceStrategy`: Runs on-device inference via `DeviceRunnerAdapter` - Validates artifact_paths and adapter are present (no default adapter — device config is required) - Push → execute → pull results flow - Two-step protocol: uses `output_data` from execute if present, falls back to pulled file paths - Returns inference results, performance metrics, and optional `ETDump` #### `DefaultCompilerAdapter`: - `compile_model` signature finalised — mirrors `to_edge_transform_and_lower_to_qnn` argument-for-argument, so per-graph lowering inputs (`compile_specs`, `dep_table`, `passes_job`, `constant_methods`) are explicit parameters rather than `extra_options` keys - Body deliberately raises `NotImplementedError`: the version this package needs is the multi-graph one (graph-name-keyed dicts, single multi-method `.pte` for weight sharing), so it lands with the strategy-level fan-out that calls it rather than being written single-graph and then replaced. Inject a custom `CompilerAdapter` for now. #### Config addition: - `CompilationInputConfig.example_inputs: Optional[Tuple[Any, ...]]` — sourced from the model via `ModelLoaderAdapter.get_example_inputs` #### Orchestrator wiring: - `genai_pipeline.py` `_run_compilation`: passes `example_inputs=model_prep_output.example_inputs` #### Unit tests: - `test_executorch_compilation_strategy.py` - `test_executorch_inference_strategy.py` - `test_compilation_input_config.py` ### PR Review Checklist - All new classes follow single responsibility (one class per file) - Yes. - All dependencies are injected via constructor with sensible defaults - Yes. - All external calls are behind injectable interfaces - Yes (adapter pattern). - Unit tests cover every public method - Yes (100% coverage on strategy impls). - Docstrings on all public classes and methods - Yes. - Type annotations on all function signatures - Yes. - Logging follows the strategy in the LLD - Yes (info on entry/exit, debug per step). ### Related PRs - PR 1: Core data model, engine routing & exceptions: #20409 - PR 2: Strategy interfaces & stage wrappers: #20795 - PR 3: Pipeline orchestrator: #21149 - PR 4: Adapter interfaces, default implementations & dataset providers: #21751 - PR 5: Model preparation & quantization strategies: #21899 - PR 6: Compilation & inference strategy implementations: this pr. - PR 7: Integration & E2E tests: pending. ## Test plan ### Run only tests added in this PR: ``` python -m pytest \ backends/qualcomm/genai_pipeline/tests/strategies/compilation/ \ backends/qualcomm/genai_pipeline/tests/strategies/inference/ \ backends/qualcomm/genai_pipeline/tests/configs/test_compilation_input_config.py -v ``` ### Run only this PR's tests with coverage: ``` python -m pytest \ backends/qualcomm/genai_pipeline/tests/strategies/compilation/ \ backends/qualcomm/genai_pipeline/tests/strategies/inference/ \ backends/qualcomm/genai_pipeline/tests/configs/test_compilation_input_config.py \ --cov=backends/qualcomm/genai_pipeline/strategies/compilation \ --cov=backends/qualcomm/genai_pipeline/strategies/inference \ --cov=backends/qualcomm/genai_pipeline/configs/compilation_input_config \ --cov-report=term-missing ``` Result: ``` Name Stmts Miss Branch BrPart Cover Missing ---------------------------------------------------------------------------------------------------------------------------------------- backends/qualcomm/genai_pipeline/strategies/compilation/compilation_strategy.py 7 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/compilation/compiler_adapter.py 11 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/compilation/executorch_compilation_strategy.py 45 0 10 0 100% backends/qualcomm/genai_pipeline/strategies/inference/device_runner_adapter.py 14 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/inference/executorch_inference_strategy.py 43 0 6 0 100% backends/qualcomm/genai_pipeline/strategies/inference/inference_strategy.py 7 0 0 0 100% ---------------------------------------------------------------------------------------------------------------------------------------- TOTAL 127 0 16 0 100% ``` ### Run all `genai_pipeline` tests: ``` python -m pytest backends/qualcomm/genai_pipeline/tests/ -v ``` ### Run all tests with coverage: ``` python -m pytest backends/qualcomm/genai_pipeline/tests/ \ --cov=backends/qualcomm/genai_pipeline \ --cov-config=backends/qualcomm/.coveragerc \ --cov-report=term-missing ``` Result: ``` Name Stmts Miss Branch BrPart Cover Missing ---------------------------------------------------------------------------------------------------------------------------------------------------- backends/qualcomm/genai_pipeline/configs/compilation_input_config.py 12 0 0 0 100% backends/qualcomm/genai_pipeline/configs/compilation_output_config.py 8 0 0 0 100% backends/qualcomm/genai_pipeline/configs/inference_input_config.py 12 0 0 0 100% backends/qualcomm/genai_pipeline/configs/inference_output_config.py 9 0 0 0 100% backends/qualcomm/genai_pipeline/configs/model_preparation_input_config.py 7 0 0 0 100% backends/qualcomm/genai_pipeline/configs/model_preparation_output_config.py 12 0 0 0 100% backends/qualcomm/genai_pipeline/configs/quantization_input_config.py 13 0 0 0 100% backends/qualcomm/genai_pipeline/configs/quantization_output_config.py 5 0 0 0 100% backends/qualcomm/genai_pipeline/datasets/calibration_data_adapter.py 5 0 0 0 100% backends/qualcomm/genai_pipeline/datasets/default_calibration_data_adapter.py 26 0 6 0 100% backends/qualcomm/genai_pipeline/datasets/default_training_data_adapter.py 14 0 2 0 100% backends/qualcomm/genai_pipeline/datasets/training_data_adapter.py 5 0 0 0 100% backends/qualcomm/genai_pipeline/engine_proxy.py 20 0 4 0 100% backends/qualcomm/genai_pipeline/exceptions.py 20 0 6 0 100% backends/qualcomm/genai_pipeline/genai_pipeline.py 99 7 12 1 93% 190-201 backends/qualcomm/genai_pipeline/pipeline_context.py 52 0 14 0 100% backends/qualcomm/genai_pipeline/pipeline_stage.py 5 0 0 0 100% backends/qualcomm/genai_pipeline/stages/compilation_stage.py 14 0 0 0 100% backends/qualcomm/genai_pipeline/stages/inference_stage.py 14 0 0 0 100% backends/qualcomm/genai_pipeline/stages/model_preparation_stage.py 14 2 0 0 86% 30, 37 backends/qualcomm/genai_pipeline/stages/quantization_stage.py 14 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/compilation/compilation_strategy.py 7 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/compilation/compiler_adapter.py 11 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/compilation/executorch_compilation_strategy.py 45 0 10 0 100% backends/qualcomm/genai_pipeline/strategies/inference/device_runner_adapter.py 14 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/inference/executorch_inference_strategy.py 43 0 6 0 100% backends/qualcomm/genai_pipeline/strategies/inference/inference_strategy.py 7 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/model_preparation/executorch_model_preparation_strategy.py 68 0 14 0 100% backends/qualcomm/genai_pipeline/strategies/model_preparation/model_loader_adapter.py 9 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/model_preparation/model_preparation_strategy.py 7 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/quantization/executorch_quantization_strategy.py 61 0 18 0 100% backends/qualcomm/genai_pipeline/strategies/quantization/quantization_strategy.py 7 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/quantization/quantizer_adapter.py 9 0 0 0 100% ---------------------------------------------------------------------------------------------------------------------------------------------------- TOTAL 668 9 92 1 99% ```
Summary
This PR adds the adapter layer that wraps external APIs (ExecuTorch, QNN SDK, HuggingFace) behind injectable Protocol interfaces for testability. Each strategy implementation (in subsequent PRs) delegates to these adapters rather than calling external APIs directly, enabling unit tests with mocked dependencies.
What's included
Adapter Protocols (6 files):
QuantizerAdapter: Protocol wrappingmake_quantizer,prepare_pt2e,calibrate,convert_pt2eCompilerAdapter: Protocol wrappingExportSessioncompilation flow +CompilationResultdataclassDeviceRunnerAdapter: Protocol wrappingSimpleADBpush/execute/pull +InferenceResultdataclassModelLoaderAdapter: Protocol wrapping HuggingFace model/tokenizer loadingCalibrationDataAdapter: Protocol for calibration dataset constructionTrainingDataAdapter: Protocol for QAT training data (yields (features, labels) pairs)Default Implementations (6 files):
DefaultQuantizerAdapter: Delegates toexport_utils.make_quantizer+torchao.quantization.pt2eDefaultCompilerAdapter: Placeholder for recipe-based compilation (depends onExportRecipe/ExportSessionAPIs not yet available). RaisesNotImplementedErrorwith guidance to inject a customCompilerAdapterusingto_edge_transform_and_lower_to_qnn.DefaultDeviceRunnerAdapter: Delegates toSimpleADBfor on-device executionDefaultModelLoaderAdapter: Delegates to HuggingFaceAutoModelForCausalLM+AutoTokenizerDefaultCalibrationDataAdapter: Random token sequences; accepts a caller-supplied dataset or DataLoader via extra_options["dataset"]DefaultTrainingDataAdapter: Pass-through for caller-supplied QAT data; raises ValueError if absent (labelled data can't be synthesized)Configuration:
.coveragercupdated to omitdefault_*_adapter.pyfiles (integration-test-only, require real SDK/hardware)__init__.pyfiles updated to export new adapter typesNew
datasets/packageDataset providers are a cross-stage concern, not a model-preparation detail — the same corpus feeds PTQ calibration during quantization and
on-device result evaluation during inference (including pre-built
.pteflows where model preparation never runs).They therefore live in a top-level
datasets/package rather than understrategies/model_preparation/.PR Review Checklist
Related PRs
Test plan
All existing tests continue to pass (no regressions).
Test Coverage
Command to run:
Result: