Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 4 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -67,8 +67,9 @@ The skill runs as a guided, three-stage workflow (the detailed algorithm lives i
2. **Plan.** Explore the object, gather prior knowledge for the tissue/context from the
knowledge stack, draft the label hierarchy and per-level vocabulary, present for review.
3. **Execute (top-down, recursive).** For each level, and within each parent group:
assemble candidate types + markers, panel-gate them, score (Wilcoxon + COSG markers;
AUCell/pyUCell signatures with a z-scored + consolidated cross-check), cross-check
assemble candidate types + markers, panel-gate them, score (Wilcoxon DE + a
`filter_rank_genes_groups` specificity pass; AUCell/pyUCell signatures with a
z-scored + consolidated cross-check), cross-check
against a tissue-matched classifier / a reference mapping (CellMapper) as a
*hypothesis*, sign off the assumptions ledger, profile ambiguous clusters before any
`Unknown`, write per-cluster evidence cards, then assign a label — or sub-cluster and
Expand Down Expand Up @@ -143,7 +144,7 @@ skills/cell-type-annotation/
│ ├── artifacts.md # working-dir layout, run_state schema, reports
│ ├── compute_environments.md # the 3 compute profiles, env detection, scheduler handoff
│ ├── tool_registry.md # canonical step → cta command → GPU/profile map
│ ├── marker_tools.md # decoupler, pyUCell, COSG, snapseed
│ ├── marker_tools.md # decoupler, pyUCell, snapseed
│ ├── automated_tools.md # CellTypist, CellMapper, scANVI/scArches, Census, what to skip
│ ├── knowledge_sources.md # MCP coverage, CL normalization, atlases, loaders
│ └── differential_expression.md # markers (Job A) vs condition DE (Job B)
Expand Down
8 changes: 4 additions & 4 deletions skills/cell-type-annotation/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -111,7 +111,7 @@ and `plan.md` in the working directory first, then continue from
- Verify whichever MCPs are present actually respond before proceeding.

**1b. Detect the compute environment (do this early — it shapes the plan).** Run
`cta detect-env` (stdlib-only — works with any python). It reports the accelerator
`cta detect-env` (no heavy imports — no scanpy/torch, so it's a fast cold start). It reports the accelerator
(CUDA / Apple-Silicon MPS / none), usable RAM (clamped to cgroup/scheduler limits), any
batch scheduler, and a **suggested profile**: `local_cpu` / `local_gpu` / `scheduler`.
**Confirm it with the user** — and you *must* ask in the one ambiguous case the probe flags:
Expand Down Expand Up @@ -383,15 +383,15 @@ first, so a deep recursion never leaves the user lost.
an L1 compartment (NPC vs Neuron), *qualify the label by compartment*
(`Ventral telencephalic NPC` vs `… neuron`) — never reuse one label across two parents.
Verify any time with `cta report write-back --dry-run` (runs
`cta_io.check_hierarchy_nesting`, refusing a non-nested hierarchy). Why the CSV is the
`cta.io.check_hierarchy_nesting`, refusing a non-nested hierarchy). Why the CSV is the
source of truth, and the full card schema: `reference/methodology.md` §4.
7. **Critic pass.** Spawn independent critic subagents (Agent tool) that read the
exported **raw tables** (not just the card), re-derive markers, fetch ≥1 external
hypothesis, and try to **refute** each label — biologically **and hierarchically**:
is the label a coherent refinement of its parent path (flag maturation/regional
mismatches like an "IPC"/"neuron" L3 under an "NPC" L1 — the kind the deterministic
nesting check can't catch)? Reconcile; revise or downgrade. (Structural nesting itself
is enforced in code — `cta_io.check_hierarchy_nesting` / `cta report write-back`.)
is enforced in code — `cta.io.check_hierarchy_nesting` / `cta report write-back`.)
8. **Report, gate & descend.** Build the level's HTML report with
`cta report build` (pass `--adata --cluster-key --viz-key`, plus
`--embedding-key` for the similarity-ordered palette, to embed the numbered
Expand Down Expand Up @@ -465,7 +465,7 @@ and **whether it needs a GPU**, `reference/tool_registry.md` is the canonical ma
| `cta reference train-model` | GPU/large-memory: train (or refine) a scANVI reference model on a labeled atlas (scvi-tools). Reference side of the train-then-map workflow. |
| `cta reference map-scarches` | GPU/large-memory: map the query onto a scANVI reference (a model you trained, or a shipped scvi-tools hub model) via scArches surgery — labels + per-cell uncertainty, in the `refmap_*` schema. |
| `cta crosscheck scimilarity` | Optional **HUMAN-only** cross-check: zero-shot per-cell labels via SCimilarity (~16–27 GB RAM, CPU, ~9 GB annotation bundle). Foundation model → additional hypothesis only, never primary; most useful on broad/multi-tissue human queries where no CellTypist/SingleR model fits. Needs the model bundle — provision it once with `cta reference fetch-scimilarity`. |
| `cta reference fetch-scimilarity` | One-time provisioning for `cta crosscheck scimilarity`: download + extract the SCimilarity model bundle. Annotation-only by default (~9 GB on disk; ~30 GB resumable download, skips the unused 32 GB cellsearch). Point `--model-path` at the resulting `<dest>/model_v1.1`. Needs network (`module load eth_proxy` on a compute node). |
| `cta reference fetch-scimilarity` | One-time provisioning for `cta crosscheck scimilarity`: download + extract the SCimilarity model bundle. Annotation-only by default (~9 GB on disk; ~30 GB resumable download, skips the unused 32 GB cellsearch). Point `--model-path` at the resulting `<dest>/model_v1.1`. Needs network (ensure outbound internet on a firewalled compute node). |
| `cta reference fetch-census` | Optional: pull a capped (or, on a large-memory profile, larger) tissue-matched reference subset from CELLxGENE Census (no full-atlas download) to feed `cta crosscheck reference-map` when the user has no atlas and no CellTypist model fits. Needs the `cellxgene-census` extra + network. |
| `cta crosscheck obs-pred` | Cross-check from an EXISTING in-object prediction column (e.g. a previously-mapped atlas `*_pred`) — no recompute. Prefer over `cta crosscheck reference-map` when such a column exists and fits the level. |
| `cta cluster subcluster` | Refine one/more clusters: Leiden on the anchor embedding + per-sub marker-fraction / QC / composition summary. Standardized sub-clustering (in scope). |
Expand Down
6 changes: 3 additions & 3 deletions skills/cell-type-annotation/pixi.toml
Original file line number Diff line number Diff line change
Expand Up @@ -91,7 +91,7 @@ extra-index-urls = ["https://pypi.nvidia.com"]

# scimilarity pulls hnswlib, but its PyPI wheel is the problem: the 0.8.0 cp312 wheel is built
# with AVX-512 and SIGILLs `cta crosscheck scimilarity` when SCimilarity loads its prebuilt
# hnswlib kNN index on AVX2-only CPUs (verified crashing on Euler's AMD EPYC 7763 / Intel Xeon
# hnswlib kNN index on AVX2-only CPUs (verified crashing on AMD EPYC 7763 / Intel Xeon
# E3-1284L v4 — neither has AVX-512). The fix is to source hnswlib from conda-forge instead:
# conda-forge builds it portably (baseline microarch, no `-march=native`), so the same 0.8.0
# runs on any CPU (verified: the conda-forge .so has zero AVX-512/AVX2 ops and the full
Expand All @@ -104,7 +104,7 @@ hnswlib = ">=0.8"
# `rank_genes_groups`) are built against CUDA 13 and link `libcudart.so.13`, but the rest of the
# GPU stack is CUDA 12 (cuml-cu12 / cupy-cuda12x / torch-cu12, which load libcudart.so.12). Without
# a CUDA-13 runtime present the kernel .so fails to load, rapids-singlecell sets it to None, and the
# GPU DE path silently falls back to scanpy (verified on Euler). Ship the CUDA-13 cudart runtime from
# GPU DE path silently falls back to scanpy (verified in practice). Ship the CUDA-13 cudart runtime from
# conda-forge so libcudart.so.13 lands in $ENV/lib alongside .so.12 — the two majors coexist (the
# driver supports both) and the kernel loads. This is the actual enabler of issue #38's GPU path.
cuda-cudart = "13.*"
Expand All @@ -128,7 +128,7 @@ cudf-cu12 = "*"
# profile (atlas-scale ranking in seconds vs ~minutes on scanpy/CPU). Drop-in — same wilcoxon
# method, same `uns['rank_genes_groups']` output the CPU path writes; scanpy stays the CPU default.
# The `rapids-cu12` extra pulls the FULL RAPIDS stack it imports at load (cudf/cuml/cugraph/cuvs);
# the bare cuml-cu12/cudf-cu12 above (for CellMapper's kNN) are not enough — verified on Euler that
# the bare cuml-cu12/cudf-cu12 above (for CellMapper's kNN) are not enough — verified in practice that
# `import rapids_singlecell` needs cuvs too.
rapids-singlecell = { version = "*", extras = ["rapids-cu12"] }
# rapids-singlecell 0.15.2 also imports `dask.array` at module load (rapids_singlecell/_compat.py)
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -108,7 +108,7 @@ CPU host that merely has the client tools installed. `cta detect-env` flags it;
annotation kNN; the full ~41 GB download incl. the cell-search TileDB is NOT needed). Fits a
32 GB machine — optional, and HUMAN-only (skip for non-human queries). Provision the bundle
once with `cta reference fetch-scimilarity --dest DIR` (resumable ~30 GB download, extracts
annotation-only by default; on a compute node `module load eth_proxy` first).
annotation-only by default; on a firewalled compute node ensure outbound internet first).

## Which tool at which step (and what it needs)

Expand Down
2 changes: 1 addition & 1 deletion skills/cell-type-annotation/reference/methodology.md
Original file line number Diff line number Diff line change
Expand Up @@ -272,7 +272,7 @@ hierarchy must nest, so the critic also checks each label against its **parent p
(L1→…→L_k):

- *Structural* nesting (every finer label has exactly one parent) is a deterministic
invariant enforced in **code** — `cta_io.check_hierarchy_nesting`, run by
invariant enforced in **code** — `cta.io.check_hierarchy_nesting`, run by
`cta report write-back` (and any time via `--dry-run`). Don't rely on an LLM to never miss it;
the critic need not re-derive it.
- *Semantic* path-consistency is the critic's job — flag a path that **passes** the
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -13,8 +13,9 @@
unless --keep-tarball), plus ~9 GB extracted → ~39 GB peak, settling to ~9 GB.
* with --with-cellsearch: ~41 GB extracted (only if you also want cell SEARCH).

The download is resumable (re-run the command to continue a partial download). On an Euler
COMPUTE node remember to ``module load eth_proxy`` first; a LOGIN node has internet directly.
The download is resumable (re-run the command to continue a partial download). On a firewalled
compute node, make sure outbound internet is available first (e.g. load your site's proxy
module); an internet-connected (login) node works directly.

Usage:
cta reference fetch-scimilarity --dest /path/to/models
Expand Down Expand Up @@ -67,8 +68,8 @@ def _download(url: str, tarball: Path) -> None:
except subprocess.CalledProcessError as e:
raise SystemExit(
f"ERROR: download failed (exit {e.returncode}). It is resumable — re-run the "
"same command to continue. On an Euler COMPUTE node, `module load eth_proxy` "
"first (a login node has internet directly)."
"same command to continue. On a firewalled compute node, ensure outbound "
"internet/proxy is configured first (an internet-connected node works directly)."
) from e


Expand Down
Loading