Skip to content

feat(cta): first-class negative markers + CL-ID validation - #42

Merged
Marius1311 merged 2 commits into
mainfrom
claude/cap-skill-exploration-45d38b
Jul 9, 2026
Merged

Marius1311 merged 2 commits into
mainfrom
claude/cap-skill-exploration-45d38b

Conversation

@Marius1311

Copy link
Copy Markdown
Member

Context

I explored the Cell Annotation Platform (CAP) (github.com/cellannotation, celltype.info) to establish whether any of it was useful for the skill. Conclusion: its schema (CAS) and tooling are a container/serializer with no bundled vocabulary or marker content and a heavy, near-unused dependency tree; its queryable corpus (cap-sc-client) is thin and uneven exactly in the developmental/organoid domain we care about. Neither was worth adopting.

But the comparison surfaced two genuine gaps in our own evidence format — worth doing regardless of CAP. This PR implements both.

1. Negative markers as scored, contestable evidence

Negative (expected-absent) markers are often the decisive discriminator (NPC = SOX2+/DCX⁻, T cell = CD3D+/MS4A1⁻) and a negative marker turning up expressed is falsifiable evidence against a call — exactly what the critic pass exists to act on. Previously they had nowhere scored to live.

  • Marker JSON is signed ({type: {"pos": [...], "neg": [...]}}); markers assemble emits an empty neg slot to curate (sources yield positives only, so negatives are curated from the sibling types' markers via the knowledge stack).
  • markers profile is sign-aware and writes panel_sign_<level>.csv + panel_neg_flags_<level>.csv. Sign is resolved per (type, gene), not collapsed per gene — a gene is routinely a positive marker of one type and an expected-absent marker of another, and both roles are kept. A violation = a negative marker expressed in more than NEG_MARKER_MAX_PCT of a cluster's cells. Off-panel negatives are reported UNMEASURED, never "confirmed absent".
  • report scaffold-cards renders an "Expected-absent (negative) markers" block, flagging violations with a ⚠ for the annotator and critic to weigh.

Threshold calibrated on real data

NEG_MARKER_MAX_PCT was validated on five real atlases (Immune Cell Atlas, GTEx v9, mouse pancreas, HypoMap, and HNOCA — the organoid atlas — via celltype-anno-bench on Euler). Genuine positive markers sit at ~52–68% fraction-expressing while true-negative background (ambient + shared developmental programs) reaches ~10–20% in developmental data. So:

threshold mean false-violation HNOCA (organoid) real positives cleared
0.10 (initial) ~18% 43% 0%
0.20 (chosen) ~5% 10% 0%
0.25 ~2% 4% 0%

0.20 cuts false alarms ~4× with zero cost to catching real co-expression; tunable toward 0.25 for very continuous/developmental data.

2. CL-ID validation at the write-back gate

cl_id was a free-text column that nothing checked — a typo (CL:000679) or wrong/obsolete id sailed silently into the user's AnnData. report write-back now syntactically validates every non-blank cl_id (CL/PCL:0000000), blocking by default with --allow-bad-cl-id, alongside the existing hierarchy-nesting gate. The semantic check (exists / not obsolete / matches label) is documented as an in-chat knowledgebase-MCP step, since cta commands have no MCP access.

Not adopted (deliberately)

No cas-tools dependency, no CAS-JSON export, no cap-sc-client knowledge source, no accession-based hierarchy — evidence in the session showed each to be low value / high cost for this use.

Testing

  • New tests/test_negative_markers.py (sign preservation, per-(type,gene) roles incl. the pos-elsewhere/neg-here case, violation thresholding, off-panel handling, dot-plot render, card rendering) and tests/test_cl_id_validation.py (accept/reject + dry-run block/override).
  • End-to-end verified on tiny + real data; caught and fixed a variable-shadowing bug in the dot-plot path along the way.
  • Full suite: 47 passed, ruff clean.

🤖 Generated with Claude Code

Marius1311 and others added 2 commits July 9, 2026 10:45
Two rigor upgrades surfaced by evaluating the Cell Annotation Platform (CAP)
against the skill. CAP's schema (CAS) and corpus weren't worth adopting, but the
comparison exposed two gaps in our own evidence format — implemented here.

Negative markers as scored, contestable evidence:
- Marker JSON is signed ({type: {"pos": [...], "neg": [...]}}); `markers assemble`
  now emits an empty `neg` slot to curate. `markers profile` tracks sign PER
  (type, gene) — a gene is routinely a positive marker of one type and an
  expected-absent marker of another, and both roles are kept (not collapsed to one
  role per gene). Writes panel_sign_<level>.csv and panel_neg_flags_<level>.csv
  (per cluster×type×neg-marker; a violation = marker expressed where the call says
  it should be absent → evidence against that call). Off-panel negatives are
  reported as UNMEASURED, never "confirmed absent".
- `report scaffold-cards` renders an "Expected-absent (negative) markers" block,
  flagging violations with a ⚠ for the annotator and critic pass to weigh.
- NEG_MARKER_MAX_PCT default calibrated on 5 real atlases (Immune Cell Atlas,
  GTEx v9, mouse pancreas, HypoMap, HNOCA via celltype-anno-bench): genuine
  positives sit at ~52-68% expressing while true-negative background reaches
  ~10-20% in developmental data, so 0.20 clears zero real positives while cutting
  the false-violation rate from ~18% (at 0.10) to ~5% mean — and 43%→10% on the
  HNOCA organoid atlas. Tunable toward 0.25 for very continuous data.

CL-ID validation at the write-back gate:
- `report write-back` syntactically validates any non-blank cl_id (CL/PCL:0000000),
  blocking by default with --allow-bad-cl-id, alongside the existing nesting gate.
  Semantic checks (exists / not obsolete / matches label) are documented as an
  in-chat knowledgebase-MCP step, since cta commands have no MCP access.

Docs (SKILL.md, marker_tools/knowledge_sources/artifacts) and tests updated;
full suite green (47 passed).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Expand compact subprocess arg lists and long lines to ruff-format's
canonical layout; no behavior change. Fixes the pre-commit CI hook.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@Marius1311
Marius1311 merged commit d5f55f2 into main Jul 9, 2026
2 checks passed
@Marius1311
Marius1311 deleted the claude/cap-skill-exploration-45d38b branch July 9, 2026 08:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant