Skip to content

feat(cta): single-source-of-truth method routing + pan-human Azimuth cross-check - #43

Merged
Marius1311 merged 1 commit into
mainfrom
feat/method-routing-and-azimuth
Aug 6, 2026
Merged

Marius1311 merged 1 commit into
mainfrom
feat/method-routing-and-azimuth

Conversation

@Marius1311

@Marius1311 Marius1311 commented Aug 6, 2026 •

Copy link
Copy Markdown
Member

Two coupled changes: make method selection an explicit, machine-checkable decision tree, and add pan-human Azimuth as a cross-check.

The problem

Method selection had no single home. The only codified routing was an else-cascade in SKILL.md prose branching on compute profile and reference availability; species, developmental stage and technology gated methods only in scattered prose across automated_tools.md, tool_registry.md and compute_environments.md. Nothing kept them consistent, and they weren't.

Routing is now data, not prose

src/cta/routing.py the ranking and every disqualifier, as one ordered table — the definition
cta route what Stage 3 step 4 runs: names the ONE primary cross-check, why each higher option was skipped or refused, and a caveat to carry into the cards
reference/method_selection.md the human page, generated from that module between markers
tests/test_method_selection_sync.py fails on drift
$ cta route --species human --stage adult --state normal --profile local_cpu
PRIMARY: cta crosscheck azimuth
  because: human, adult/postnatal, non-tumour, whole-transcriptome — and no matched reference above
  CAVEAT : … it CANNOT be constrained to a tissue … on immune tissue it is NOT an independent
           second opinion relative to CellTypist …

ranked evaluation:
    1. cta crosscheck obs-pred        skipped       preconditions not met
  x 2. cta reference map-scarches     disqualified  no CUDA GPU reachable from this process
    3. cta crosscheck reference-map   skipped       preconditions not met
 -> 4. cta crosscheck azimuth         primary       …
  x 5. cta crosscheck celltypist      disqualified  no tissue- AND stage-matched model — OMIT rather than force one
    6. cta crosscheck scimilarity     fallback      outranked
    7. cta reference fetch-census …   fallback      outranked

The sync test enforces: the doc equals what the module emits; every routed command exists in GROUPS and has a tool_registry.md row; exactly one primary for any context; no query is ever left with no cross-check; the disqualifiers can't be loosened; and SKILL.md does not restate the cascade.

SKILL.md's cascade is deleted, not kept as an inline "fast path" — keeping both is exactly the two-definitions problem this removes.

What the ranking is, and is not

The ordering encodes a preference: a curated, tissue-matched reference outranks a generic one, then cost and what each tool returns beyond a label. It is not a claim that any tool annotates better than another, and nothing here should be read as a published benchmark result. The disqualifiers come from each tool's documented scope and structural limits — not from our own scoring.

Nine inconsistencies the tree exposed

Each is a place where following the docs literally led to the wrong action:

  1. "Run exactly one cross-check" (SKILL.md) vs a recommended 2–3 method mini-consensus (automated_tools.md) — a direct contradiction. Now tie-break-only.
  2. SCimilarity was unreachable: filed under a "GPU / large-memory tier" heading, listed as profile all in the registry, and not a branch in the cascade at all. The registry advertised a tool the executable loop could not reach.
  3. CellTypist was "tissue-gated" when stage matters equally — a stage-mismatched model collapses rather than degrading gracefully, so the rule is omit it, not force one.
  4. Species appeared nowhere in the cascade, despite gating several methods.
  5. Nothing said a pan-transcriptome model is disqualified on a targeted panel. Now the tree's first branch, and cta crosscheck azimuth refuses below --min-panel-overlap rather than zero-filling most of its input into confident labels.
  6. "Foundation models rarely beat scANVI" vs "the strongest automated option" — reworded to "rarely beat a well-matched one, so they are the fallback, never the default."
  7. obs-pred silently inherits its producer's disqualifiers — now a stated caveat.
  8. mean_conf means something different in every refmap_* table (scArches posterior / CellMapper confidence / within-query distance inversion / calibrated probability) while the cards present them identically. Documented.
  9. Known gap, deliberately not fixed: SingleR is recommended in automated_tools.md and installed in pixi.toml, but has no command, no registry row and no route — while tool_registry.md calls itself "the single source of truth for tool selection". Filed as SingleR is recommended and installed but unroutable — wrap it or demote it #45.

cta crosscheck azimuth

A supervised classifier (8 chained hierarchical heads over a fixed 5,055-gene panel; ~9.7M cells, 23 human tissues, cancer excluded), CPU-only at ~50–60k cells/min, 84 MB model via cta reference fetch-azimuth.

Ranked 3rd — below a matched scANVI model or the user's own curated atlas, above CellTypist / Census / SCimilarity — because it is cheap to consult and returns more than a label.

Its disqualifiers follow the authors' own stated scope: the model "was not designed to capture developmental stages, disease-specific states, or cancer programs", and is human-only. The targeted-panel disqualifier is structural — panhumanpy zero-fills panel genes the query lacks.

What it uniquely provides among the cross-checks here:

  • a Cell Ontology id per cluster (with SKOS match strength)
  • a calibrated confidence (temperature-scaled — the only mean_conf readable as a probability)
  • trained Unassigned and Doublet like cell classes

The last two land in azimuth_state_<level>.csv and are documented in methodology.md §6 as a second, independent doublet / low-quality corroborator — until now the only in-object doublet signal was a pre-existing Scrublet/scDblFinder column. This adds one without bringing doublet detection into scope. frac_inconsistent_hierarchy doubles as an out-of-distribution alarm. These corroborate a state; they never set one.

Two limits carried in the routing caveat so they reach the evidence cards: it cannot be constrained to a tissue (on a single-organ query it may return an out-of-organ type), and its training labels on lymph node/spleen derive from the CellTypist organ atlases, so there it is not an independent second opinion.

Guards. Only counts or lognorm_cp10k are accepted — panhumanpy reads .X only and infers normalization by sniffing integer-ness, so the canonical log-norm X (the user's own, not necessarily CP10K) would be silently mis-scaled. Ensembl var_names are refused (upstream #39: it returns nonsense rather than failing). Panel overlap is gated.

Environment

A new azimuth pixi env layering panhumanpy on the normal numpy>=2 base via dependency-overrides — not a fully isolated env. panhumanpy's tensorflow==2.17 / scikit-learn==1.6.0 exact pins are not load-bearing: verified on 2,700 cells that the same v1 model gives identical labels under TF 2.17/numpy 1.26/sklearn 1.6 and TF 2.21/numpy 2.4/sklearn 1.9, confidences differing by ≤3.6e-06. Keras resolves to the same 3.15.1 in both — and Keras, not TF, is what loads the .keras artifacts. Filed as satijalab/panhumanpy#49. Kept out of default because it drags in TensorFlow for one optional cross-check.

Expect a large pixi.lock diff; CI only installs default, so runtime cost is zero.

Verification

  • pixi run lint clean; 70 passed, 1 skipped in default
  • the mirrored-constants drift guard passes in the azimuth env (skips in default)
  • cta route validated across ~10 scenarios (mouse / organoid / tumour / MERSCOPE panel / GPU+scANVI / low-RAM)
  • cta crosscheck azimuth run end-to-end on a canonicalized object: all three tables written, both guards fire with actionable messages

Not yet done: a full blinded skill run with the new Stage-3 step 4 — filed as #44.

🤖 Generated with Claude Code

…cross-check

Method selection had no single home. The only codified routing was an else-cascade
in SKILL.md prose (compute profile + reference availability); species, stage and
technology gated methods only in scattered prose across three reference docs. So
the docs, the tool registry and the actual behaviour could drift — and did.

## Routing is now data, not prose

* `src/cta/routing.py` — the ranking and every disqualifier, as one ordered table.
* `cta route` — what Stage 3 step 4 runs. Names the ONE primary cross-check and
  why each higher-ranked option was skipped or refused, plus a caveat to carry
  into the evidence cards.
* `reference/method_selection.md` — the human-readable page, GENERATED from that
  module between markers (`cta route --emit-markdown`).
* `tests/test_method_selection_sync.py` fails if the doc drifts, if a route names
  a command the CLI lacks or the tool registry omits, if two primaries are ever
  chosen, if any query is left with no cross-check, or if a disqualifier is
  loosened. A test also fails if SKILL.md restates the cascade inline.

SKILL.md's cascade is DELETED rather than kept as a "fast path" — keeping both is
exactly the two-definitions problem this change exists to remove.

The ranking encodes a PREFERENCE, not a performance claim: a curated,
tissue-matched reference outranks a generic one, then cost and what each tool
returns beyond a label. The disqualifiers come from each tool's documented scope
and structural limits, not from our own scoring.

## Nine inconsistencies the tree exposed, fixed

Writing the routing down surfaced places where following the docs literally led
to the wrong action:

* "run exactly one cross-check" (SKILL.md) vs a recommended 2-3 method
  mini-consensus (automated_tools.md) — the latter is now tie-break-only.
* SCimilarity was filed under a "GPU / large-memory tier" heading, listed as
  profile `all` in the registry, and was not a branch in the cascade at all — so
  the executable loop could not reach a tool the registry advertised. Pan-body
  models now have their own all-profiles section and are reachable via `cta route`.
* CellTypist was "tissue-gated" when stage matters equally — a stage-mismatched
  model collapses rather than degrading gracefully, so the rule is to OMIT it.
* Species appeared nowhere in the cascade despite gating several methods.
* No doc said a pan-transcriptome model is DISQUALIFIED on a targeted panel —
  now the tree's first branch, and `cta crosscheck azimuth` refuses below
  --min-panel-overlap rather than zero-filling into confident garbage.
* `mean_conf` means something different in every refmap_* table; documented.
* obs-pred inherits its producer's disqualifiers; stated as a caveat.

Left as a known gap: SingleR is recommended in automated_tools.md and installed
in pixi.toml but has no command, no registry row and no route.

## `cta crosscheck azimuth`

Pan-human Azimuth: a supervised classifier (8 chained hierarchical heads over a
fixed 5,055-gene panel; ~9.7M cells, 23 human tissues) — CPU-only, ~50-60k
cells/min, 84 MB model via `cta reference fetch-azimuth`.

Ranked 3rd, below a matched scANVI model or the user's own atlas and above
CellTypist/Census/SCimilarity, because it is cheap to consult and returns more
than a label. Its disqualifiers follow the authors' stated scope — the model
"was not designed to capture developmental stages, disease-specific states, or
cancer programs", and is human-only — plus the structural fact that panhumanpy
zero-fills missing panel genes, which rules out targeted imaging panels.

What it uniquely provides: a Cell Ontology id per cluster, a CALIBRATED
confidence, and trained `Unassigned` / `Doublet like cell` classes. Those land in
`azimuth_state_<level>.csv` and are documented in methodology.md §6 as a SECOND
independent doublet/low-quality corroborator — until now the only in-object
doublet signal was a pre-existing Scrublet/scDblFinder column.
`frac_inconsistent_hierarchy` doubles as an out-of-distribution alarm.

Two limits carried in the routing caveat: it CANNOT be constrained to a tissue
(so on a single-organ query it may return an out-of-organ type), and its training
labels on lymph node/spleen derive from the CellTypist organ atlases, so there it
is not an independent second opinion.

Guards: only `counts` or `lognorm_cp10k` are accepted (panhumanpy reads .X only
and infers normalization by sniffing integer-ness, so the canonical log-norm `X`
— the user's own, not necessarily CP10K — would be silently mis-scaled); Ensembl
var_names are refused (upstream #39: panhumanpy returns nonsense rather than
failing); panel overlap is gated.

## Environment

New `azimuth` pixi environment layering panhumanpy on the normal numpy>=2 base
via `dependency-overrides`, NOT a fully isolated env. panhumanpy's
`tensorflow==2.17` / `scikit-learn==1.6.0` exact pins are not load-bearing:
verified on 2,700 cells that the same v1 model gives identical labels under
TF 2.17/numpy 1.26/sklearn 1.6 and TF 2.21/numpy 2.4/sklearn 1.9, with
confidences differing by <=3.6e-06 (Keras resolves to the same 3.15.1 in both,
and Keras — not TF — is what loads the `.keras` artifacts). Filed upstream as
satijalab/panhumanpy#49. Kept out of `default` because it drags in TensorFlow
for one optional cross-check.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@Marius1311
Marius1311 force-pushed the feat/method-routing-and-azimuth branch from 72db305 to 925156d Compare August 6, 2026 13:37
@Marius1311
Marius1311 merged commit b438119 into main Aug 6, 2026
2 checks passed
@Marius1311
Marius1311 deleted the feat/method-routing-and-azimuth branch August 6, 2026 14:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant