feat(cta): single-source-of-truth method routing + pan-human Azimuth cross-check - #43
Merged
Merged
Conversation
This was referenced Aug 6, 2026
Marius1311
force-pushed
the
feat/method-routing-and-azimuth
branch
from
August 6, 2026 13:30
5e2ad9c to
72db305
Compare
…cross-check Method selection had no single home. The only codified routing was an else-cascade in SKILL.md prose (compute profile + reference availability); species, stage and technology gated methods only in scattered prose across three reference docs. So the docs, the tool registry and the actual behaviour could drift — and did. ## Routing is now data, not prose * `src/cta/routing.py` — the ranking and every disqualifier, as one ordered table. * `cta route` — what Stage 3 step 4 runs. Names the ONE primary cross-check and why each higher-ranked option was skipped or refused, plus a caveat to carry into the evidence cards. * `reference/method_selection.md` — the human-readable page, GENERATED from that module between markers (`cta route --emit-markdown`). * `tests/test_method_selection_sync.py` fails if the doc drifts, if a route names a command the CLI lacks or the tool registry omits, if two primaries are ever chosen, if any query is left with no cross-check, or if a disqualifier is loosened. A test also fails if SKILL.md restates the cascade inline. SKILL.md's cascade is DELETED rather than kept as a "fast path" — keeping both is exactly the two-definitions problem this change exists to remove. The ranking encodes a PREFERENCE, not a performance claim: a curated, tissue-matched reference outranks a generic one, then cost and what each tool returns beyond a label. The disqualifiers come from each tool's documented scope and structural limits, not from our own scoring. ## Nine inconsistencies the tree exposed, fixed Writing the routing down surfaced places where following the docs literally led to the wrong action: * "run exactly one cross-check" (SKILL.md) vs a recommended 2-3 method mini-consensus (automated_tools.md) — the latter is now tie-break-only. * SCimilarity was filed under a "GPU / large-memory tier" heading, listed as profile `all` in the registry, and was not a branch in the cascade at all — so the executable loop could not reach a tool the registry advertised. Pan-body models now have their own all-profiles section and are reachable via `cta route`. * CellTypist was "tissue-gated" when stage matters equally — a stage-mismatched model collapses rather than degrading gracefully, so the rule is to OMIT it. * Species appeared nowhere in the cascade despite gating several methods. * No doc said a pan-transcriptome model is DISQUALIFIED on a targeted panel — now the tree's first branch, and `cta crosscheck azimuth` refuses below --min-panel-overlap rather than zero-filling into confident garbage. * `mean_conf` means something different in every refmap_* table; documented. * obs-pred inherits its producer's disqualifiers; stated as a caveat. Left as a known gap: SingleR is recommended in automated_tools.md and installed in pixi.toml but has no command, no registry row and no route. ## `cta crosscheck azimuth` Pan-human Azimuth: a supervised classifier (8 chained hierarchical heads over a fixed 5,055-gene panel; ~9.7M cells, 23 human tissues) — CPU-only, ~50-60k cells/min, 84 MB model via `cta reference fetch-azimuth`. Ranked 3rd, below a matched scANVI model or the user's own atlas and above CellTypist/Census/SCimilarity, because it is cheap to consult and returns more than a label. Its disqualifiers follow the authors' stated scope — the model "was not designed to capture developmental stages, disease-specific states, or cancer programs", and is human-only — plus the structural fact that panhumanpy zero-fills missing panel genes, which rules out targeted imaging panels. What it uniquely provides: a Cell Ontology id per cluster, a CALIBRATED confidence, and trained `Unassigned` / `Doublet like cell` classes. Those land in `azimuth_state_<level>.csv` and are documented in methodology.md §6 as a SECOND independent doublet/low-quality corroborator — until now the only in-object doublet signal was a pre-existing Scrublet/scDblFinder column. `frac_inconsistent_hierarchy` doubles as an out-of-distribution alarm. Two limits carried in the routing caveat: it CANNOT be constrained to a tissue (so on a single-organ query it may return an out-of-organ type), and its training labels on lymph node/spleen derive from the CellTypist organ atlases, so there it is not an independent second opinion. Guards: only `counts` or `lognorm_cp10k` are accepted (panhumanpy reads .X only and infers normalization by sniffing integer-ness, so the canonical log-norm `X` — the user's own, not necessarily CP10K — would be silently mis-scaled); Ensembl var_names are refused (upstream #39: panhumanpy returns nonsense rather than failing); panel overlap is gated. ## Environment New `azimuth` pixi environment layering panhumanpy on the normal numpy>=2 base via `dependency-overrides`, NOT a fully isolated env. panhumanpy's `tensorflow==2.17` / `scikit-learn==1.6.0` exact pins are not load-bearing: verified on 2,700 cells that the same v1 model gives identical labels under TF 2.17/numpy 1.26/sklearn 1.6 and TF 2.21/numpy 2.4/sklearn 1.9, with confidences differing by <=3.6e-06 (Keras resolves to the same 3.15.1 in both, and Keras — not TF — is what loads the `.keras` artifacts). Filed upstream as satijalab/panhumanpy#49. Kept out of `default` because it drags in TensorFlow for one optional cross-check. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Marius1311
force-pushed
the
feat/method-routing-and-azimuth
branch
from
August 6, 2026 13:37
72db305 to
925156d
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Two coupled changes: make method selection an explicit, machine-checkable decision tree, and add pan-human Azimuth as a cross-check.
The problem
Method selection had no single home. The only codified routing was an
else-cascade in SKILL.md prose branching on compute profile and reference availability; species, developmental stage and technology gated methods only in scattered prose acrossautomated_tools.md,tool_registry.mdandcompute_environments.md. Nothing kept them consistent, and they weren't.Routing is now data, not prose
src/cta/routing.pycta routereference/method_selection.mdtests/test_method_selection_sync.pyThe sync test enforces: the doc equals what the module emits; every routed command exists in
GROUPSand has atool_registry.mdrow; exactly one primary for any context; no query is ever left with no cross-check; the disqualifiers can't be loosened; and SKILL.md does not restate the cascade.SKILL.md's cascade is deleted, not kept as an inline "fast path" — keeping both is exactly the two-definitions problem this removes.
What the ranking is, and is not
The ordering encodes a preference: a curated, tissue-matched reference outranks a generic one, then cost and what each tool returns beyond a label. It is not a claim that any tool annotates better than another, and nothing here should be read as a published benchmark result. The disqualifiers come from each tool's documented scope and structural limits — not from our own scoring.
Nine inconsistencies the tree exposed
Each is a place where following the docs literally led to the wrong action:
automated_tools.md) — a direct contradiction. Now tie-break-only.allin the registry, and not a branch in the cascade at all. The registry advertised a tool the executable loop could not reach.cta crosscheck azimuthrefuses below--min-panel-overlaprather than zero-filling most of its input into confident labels.obs-predsilently inherits its producer's disqualifiers — now a stated caveat.mean_confmeans something different in everyrefmap_*table (scArches posterior / CellMapper confidence / within-query distance inversion / calibrated probability) while the cards present them identically. Documented.automated_tools.mdand installed inpixi.toml, but has no command, no registry row and no route — whiletool_registry.mdcalls itself "the single source of truth for tool selection". Filed as SingleR is recommended and installed but unroutable — wrap it or demote it #45.cta crosscheck azimuthA supervised classifier (8 chained hierarchical heads over a fixed 5,055-gene panel; ~9.7M cells, 23 human tissues, cancer excluded), CPU-only at ~50–60k cells/min, 84 MB model via
cta reference fetch-azimuth.Ranked 3rd — below a matched scANVI model or the user's own curated atlas, above CellTypist / Census / SCimilarity — because it is cheap to consult and returns more than a label.
Its disqualifiers follow the authors' own stated scope: the model "was not designed to capture developmental stages, disease-specific states, or cancer programs", and is human-only. The targeted-panel disqualifier is structural —
panhumanpyzero-fills panel genes the query lacks.What it uniquely provides among the cross-checks here:
mean_confreadable as a probability)UnassignedandDoublet like cellclassesThe last two land in
azimuth_state_<level>.csvand are documented inmethodology.md§6 as a second, independent doublet / low-quality corroborator — until now the only in-object doublet signal was a pre-existing Scrublet/scDblFinder column. This adds one without bringing doublet detection into scope.frac_inconsistent_hierarchydoubles as an out-of-distribution alarm. These corroborate a state; they never set one.Two limits carried in the routing caveat so they reach the evidence cards: it cannot be constrained to a tissue (on a single-organ query it may return an out-of-organ type), and its training labels on lymph node/spleen derive from the CellTypist organ atlases, so there it is not an independent second opinion.
Guards. Only
countsorlognorm_cp10kare accepted — panhumanpy reads.Xonly and infers normalization by sniffing integer-ness, so the canonical log-normX(the user's own, not necessarily CP10K) would be silently mis-scaled. Ensemblvar_namesare refused (upstream #39: it returns nonsense rather than failing). Panel overlap is gated.Environment
A new
azimuthpixi env layering panhumanpy on the normalnumpy>=2base viadependency-overrides— not a fully isolated env. panhumanpy'stensorflow==2.17/scikit-learn==1.6.0exact pins are not load-bearing: verified on 2,700 cells that the same v1 model gives identical labels under TF 2.17/numpy 1.26/sklearn 1.6 and TF 2.21/numpy 2.4/sklearn 1.9, confidences differing by ≤3.6e-06. Keras resolves to the same 3.15.1 in both — and Keras, not TF, is what loads the.kerasartifacts. Filed as satijalab/panhumanpy#49. Kept out ofdefaultbecause it drags in TensorFlow for one optional cross-check.Expect a large
pixi.lockdiff; CI only installsdefault, so runtime cost is zero.Verification
pixi run lintclean; 70 passed, 1 skipped indefaultazimuthenv (skips indefault)cta routevalidated across ~10 scenarios (mouse / organoid / tumour / MERSCOPE panel / GPU+scANVI / low-RAM)cta crosscheck azimuthrun end-to-end on a canonicalized object: all three tables written, both guards fire with actionable messagesNot yet done: a full blinded skill run with the new Stage-3 step 4 — filed as #44.
🤖 Generated with Claude Code