Qualcomm AI Engine Direct - [GenAI Pipeline] PR5: Model preparation & quantization strategy implementations - #21899
Conversation
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21899
Note: Links to docs will display an error until the docs builds have been completed. This comment was automatically generated by Dr. CI and updates every 15 minutes. |
|
@pytorchbot label "release notes: qualcomm" |
|
@claude review this PR also review this quant_dtype is a required positional here, but the strategy omits it on the default path expecting export_utils's use_8a8w default — so ExecuTorchQuantizationStrategy() with defaults raises TypeError. Could we give quant_dtype a default here (and create_autospec the adapter in tests)? The bare MagicMock accepts any signature, so this path currently passes CI but would crash in production. |
|
Claude finished @psiddh's task in 2m 6s —— View job PR Review: PR5 — Model preparation & quantization strategies
1. Confirmed:
|
|
Accepting to unblock, @qti-horodnic can you take a look at #1 and possibly #2 (its ok if you want to address in upcoming PRs) |
… quantization strategy implementations
|
Thanks @psiddh, I've addressed the 2 comments. Should be good to merge |
Thanks again for the refactor , it is really looking solid |
…ence strategy implementations (#22284) ## Summary This PR implements the __compilation__ and __inference__ strategy implementations, completing the strategy layer. Each strategy delegates to injectable adapter interfaces for testability. ### What's included #### Strategy implementations: - `ExecuTorchCompilationStrategy`: Compiles model to .pte artifacts via `CompilerAdapter` - Validates model, example_inputs, soc_model, and backend_type are present - Passes `example_inputs` explicitly to the adapter (not via `extra_options`) — mirrors the PR5 fix for the quantization stage - Delegates to adapter with example_inputs, compile specs, artifact dir, soc_model, backend_type - Filters `context.extra_options` to a compilation-relevant allow-list - Returns artifact paths and optional `ETRecord` - `ExecuTorchInferenceStrategy`: Runs on-device inference via `DeviceRunnerAdapter` - Validates artifact_paths and adapter are present (no default adapter — device config is required) - Push → execute → pull results flow - Two-step protocol: uses `output_data` from execute if present, falls back to pulled file paths - Returns inference results, performance metrics, and optional `ETDump` #### `DefaultCompilerAdapter`: - `compile_model` signature finalised — mirrors `to_edge_transform_and_lower_to_qnn` argument-for-argument, so per-graph lowering inputs (`compile_specs`, `dep_table`, `passes_job`, `constant_methods`) are explicit parameters rather than `extra_options` keys - Body deliberately raises `NotImplementedError`: the version this package needs is the multi-graph one (graph-name-keyed dicts, single multi-method `.pte` for weight sharing), so it lands with the strategy-level fan-out that calls it rather than being written single-graph and then replaced. Inject a custom `CompilerAdapter` for now. #### Config addition: - `CompilationInputConfig.example_inputs: Optional[Tuple[Any, ...]]` — sourced from the model via `ModelLoaderAdapter.get_example_inputs` #### Orchestrator wiring: - `genai_pipeline.py` `_run_compilation`: passes `example_inputs=model_prep_output.example_inputs` #### Unit tests: - `test_executorch_compilation_strategy.py` - `test_executorch_inference_strategy.py` - `test_compilation_input_config.py` ### PR Review Checklist - All new classes follow single responsibility (one class per file) - Yes. - All dependencies are injected via constructor with sensible defaults - Yes. - All external calls are behind injectable interfaces - Yes (adapter pattern). - Unit tests cover every public method - Yes (100% coverage on strategy impls). - Docstrings on all public classes and methods - Yes. - Type annotations on all function signatures - Yes. - Logging follows the strategy in the LLD - Yes (info on entry/exit, debug per step). ### Related PRs - PR 1: Core data model, engine routing & exceptions: #20409 - PR 2: Strategy interfaces & stage wrappers: #20795 - PR 3: Pipeline orchestrator: #21149 - PR 4: Adapter interfaces, default implementations & dataset providers: #21751 - PR 5: Model preparation & quantization strategies: #21899 - PR 6: Compilation & inference strategy implementations: this pr. - PR 7: Integration & E2E tests: pending. ## Test plan ### Run only tests added in this PR: ``` python -m pytest \ backends/qualcomm/genai_pipeline/tests/strategies/compilation/ \ backends/qualcomm/genai_pipeline/tests/strategies/inference/ \ backends/qualcomm/genai_pipeline/tests/configs/test_compilation_input_config.py -v ``` ### Run only this PR's tests with coverage: ``` python -m pytest \ backends/qualcomm/genai_pipeline/tests/strategies/compilation/ \ backends/qualcomm/genai_pipeline/tests/strategies/inference/ \ backends/qualcomm/genai_pipeline/tests/configs/test_compilation_input_config.py \ --cov=backends/qualcomm/genai_pipeline/strategies/compilation \ --cov=backends/qualcomm/genai_pipeline/strategies/inference \ --cov=backends/qualcomm/genai_pipeline/configs/compilation_input_config \ --cov-report=term-missing ``` Result: ``` Name Stmts Miss Branch BrPart Cover Missing ---------------------------------------------------------------------------------------------------------------------------------------- backends/qualcomm/genai_pipeline/strategies/compilation/compilation_strategy.py 7 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/compilation/compiler_adapter.py 11 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/compilation/executorch_compilation_strategy.py 45 0 10 0 100% backends/qualcomm/genai_pipeline/strategies/inference/device_runner_adapter.py 14 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/inference/executorch_inference_strategy.py 43 0 6 0 100% backends/qualcomm/genai_pipeline/strategies/inference/inference_strategy.py 7 0 0 0 100% ---------------------------------------------------------------------------------------------------------------------------------------- TOTAL 127 0 16 0 100% ``` ### Run all `genai_pipeline` tests: ``` python -m pytest backends/qualcomm/genai_pipeline/tests/ -v ``` ### Run all tests with coverage: ``` python -m pytest backends/qualcomm/genai_pipeline/tests/ \ --cov=backends/qualcomm/genai_pipeline \ --cov-config=backends/qualcomm/.coveragerc \ --cov-report=term-missing ``` Result: ``` Name Stmts Miss Branch BrPart Cover Missing ---------------------------------------------------------------------------------------------------------------------------------------------------- backends/qualcomm/genai_pipeline/configs/compilation_input_config.py 12 0 0 0 100% backends/qualcomm/genai_pipeline/configs/compilation_output_config.py 8 0 0 0 100% backends/qualcomm/genai_pipeline/configs/inference_input_config.py 12 0 0 0 100% backends/qualcomm/genai_pipeline/configs/inference_output_config.py 9 0 0 0 100% backends/qualcomm/genai_pipeline/configs/model_preparation_input_config.py 7 0 0 0 100% backends/qualcomm/genai_pipeline/configs/model_preparation_output_config.py 12 0 0 0 100% backends/qualcomm/genai_pipeline/configs/quantization_input_config.py 13 0 0 0 100% backends/qualcomm/genai_pipeline/configs/quantization_output_config.py 5 0 0 0 100% backends/qualcomm/genai_pipeline/datasets/calibration_data_adapter.py 5 0 0 0 100% backends/qualcomm/genai_pipeline/datasets/default_calibration_data_adapter.py 26 0 6 0 100% backends/qualcomm/genai_pipeline/datasets/default_training_data_adapter.py 14 0 2 0 100% backends/qualcomm/genai_pipeline/datasets/training_data_adapter.py 5 0 0 0 100% backends/qualcomm/genai_pipeline/engine_proxy.py 20 0 4 0 100% backends/qualcomm/genai_pipeline/exceptions.py 20 0 6 0 100% backends/qualcomm/genai_pipeline/genai_pipeline.py 99 7 12 1 93% 190-201 backends/qualcomm/genai_pipeline/pipeline_context.py 52 0 14 0 100% backends/qualcomm/genai_pipeline/pipeline_stage.py 5 0 0 0 100% backends/qualcomm/genai_pipeline/stages/compilation_stage.py 14 0 0 0 100% backends/qualcomm/genai_pipeline/stages/inference_stage.py 14 0 0 0 100% backends/qualcomm/genai_pipeline/stages/model_preparation_stage.py 14 2 0 0 86% 30, 37 backends/qualcomm/genai_pipeline/stages/quantization_stage.py 14 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/compilation/compilation_strategy.py 7 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/compilation/compiler_adapter.py 11 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/compilation/executorch_compilation_strategy.py 45 0 10 0 100% backends/qualcomm/genai_pipeline/strategies/inference/device_runner_adapter.py 14 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/inference/executorch_inference_strategy.py 43 0 6 0 100% backends/qualcomm/genai_pipeline/strategies/inference/inference_strategy.py 7 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/model_preparation/executorch_model_preparation_strategy.py 68 0 14 0 100% backends/qualcomm/genai_pipeline/strategies/model_preparation/model_loader_adapter.py 9 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/model_preparation/model_preparation_strategy.py 7 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/quantization/executorch_quantization_strategy.py 61 0 18 0 100% backends/qualcomm/genai_pipeline/strategies/quantization/quantization_strategy.py 7 0 0 0 100% backends/qualcomm/genai_pipeline/strategies/quantization/quantizer_adapter.py 9 0 0 0 100% ---------------------------------------------------------------------------------------------------------------------------------------------------- TOTAL 668 9 92 1 99% ```
Summary
This PR implements the model preparation and quantization strategy implementations, replacing the
NotImplementedErrorstubs with real logic. Each strategy delegates to injectable adapter interfaces (from PR4) for testability.What's included
Strategy implementations (2 files + 1
__init__fix):ExecuTorchModelPreparationStrategy: 5-step flowload_model→load_tokenizer(viaModelLoaderAdapter)generate_calibration_data(via separately-injectableCalibrationDataAdapter)extra_optionsfallback)model_name,soc_modelrequired)ExecuTorchQuantizationStrategy: Full PT2E single-graph pipeline viaQuantizerAdapterquant_dtype,quant_recipe, and per-channel options viaextra_optionsIterableas calibration data (lists, DataLoaders, generators)training_datais provided (QAT deferred)strategies/model_preparation/__init__.py: adds missingExecuTorchModelPreparationStrategyimport to__all__Unit tests:
test_executorch_model_preparation_strategy.pytest_executorch_quantization_strategy.pytest_default_model_preparation_adapter.pytest_default_model_preparation_adapter.pyPR Review Checklist
Related PRs
Test plan
Run only tests added in this PR:
Run only this PR's tests with coverage:
Result:
Run all
genai_pipelinetests:Run all
genai_pipelinetests with coverage:Result: