feat(cta): first-class negative markers + CL-ID validation - #42
Merged
Merged
Conversation
Two rigor upgrades surfaced by evaluating the Cell Annotation Platform (CAP)
against the skill. CAP's schema (CAS) and corpus weren't worth adopting, but the
comparison exposed two gaps in our own evidence format — implemented here.
Negative markers as scored, contestable evidence:
- Marker JSON is signed ({type: {"pos": [...], "neg": [...]}}); `markers assemble`
now emits an empty `neg` slot to curate. `markers profile` tracks sign PER
(type, gene) — a gene is routinely a positive marker of one type and an
expected-absent marker of another, and both roles are kept (not collapsed to one
role per gene). Writes panel_sign_<level>.csv and panel_neg_flags_<level>.csv
(per cluster×type×neg-marker; a violation = marker expressed where the call says
it should be absent → evidence against that call). Off-panel negatives are
reported as UNMEASURED, never "confirmed absent".
- `report scaffold-cards` renders an "Expected-absent (negative) markers" block,
flagging violations with a ⚠ for the annotator and critic pass to weigh.
- NEG_MARKER_MAX_PCT default calibrated on 5 real atlases (Immune Cell Atlas,
GTEx v9, mouse pancreas, HypoMap, HNOCA via celltype-anno-bench): genuine
positives sit at ~52-68% expressing while true-negative background reaches
~10-20% in developmental data, so 0.20 clears zero real positives while cutting
the false-violation rate from ~18% (at 0.10) to ~5% mean — and 43%→10% on the
HNOCA organoid atlas. Tunable toward 0.25 for very continuous data.
CL-ID validation at the write-back gate:
- `report write-back` syntactically validates any non-blank cl_id (CL/PCL:0000000),
blocking by default with --allow-bad-cl-id, alongside the existing nesting gate.
Semantic checks (exists / not obsolete / matches label) are documented as an
in-chat knowledgebase-MCP step, since cta commands have no MCP access.
Docs (SKILL.md, marker_tools/knowledge_sources/artifacts) and tests updated;
full suite green (47 passed).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Expand compact subprocess arg lists and long lines to ruff-format's canonical layout; no behavior change. Fixes the pre-commit CI hook. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Context
I explored the Cell Annotation Platform (CAP) (
github.com/cellannotation,celltype.info) to establish whether any of it was useful for the skill. Conclusion: its schema (CAS) and tooling are a container/serializer with no bundled vocabulary or marker content and a heavy, near-unused dependency tree; its queryable corpus (cap-sc-client) is thin and uneven exactly in the developmental/organoid domain we care about. Neither was worth adopting.But the comparison surfaced two genuine gaps in our own evidence format — worth doing regardless of CAP. This PR implements both.
1. Negative markers as scored, contestable evidence
Negative (expected-absent) markers are often the decisive discriminator (
NPC = SOX2+/DCX⁻,T cell = CD3D+/MS4A1⁻) and a negative marker turning up expressed is falsifiable evidence against a call — exactly what the critic pass exists to act on. Previously they had nowhere scored to live.{type: {"pos": [...], "neg": [...]}});markers assembleemits an emptynegslot to curate (sources yield positives only, so negatives are curated from the sibling types' markers via the knowledge stack).markers profileis sign-aware and writespanel_sign_<level>.csv+panel_neg_flags_<level>.csv. Sign is resolved per(type, gene), not collapsed per gene — a gene is routinely a positive marker of one type and an expected-absent marker of another, and both roles are kept. A violation = a negative marker expressed in more thanNEG_MARKER_MAX_PCTof a cluster's cells. Off-panel negatives are reported UNMEASURED, never "confirmed absent".report scaffold-cardsrenders an "Expected-absent (negative) markers" block, flagging violations with a ⚠ for the annotator and critic to weigh.Threshold calibrated on real data
NEG_MARKER_MAX_PCTwas validated on five real atlases (Immune Cell Atlas, GTEx v9, mouse pancreas, HypoMap, and HNOCA — the organoid atlas — viacelltype-anno-benchon Euler). Genuine positive markers sit at ~52–68% fraction-expressing while true-negative background (ambient + shared developmental programs) reaches ~10–20% in developmental data. So:0.20 cuts false alarms ~4× with zero cost to catching real co-expression; tunable toward 0.25 for very continuous/developmental data.
2. CL-ID validation at the write-back gate
cl_idwas a free-text column that nothing checked — a typo (CL:000679) or wrong/obsolete id sailed silently into the user's AnnData.report write-backnow syntactically validates every non-blankcl_id(CL/PCL:0000000), blocking by default with--allow-bad-cl-id, alongside the existing hierarchy-nesting gate. The semantic check (exists / not obsolete / matches label) is documented as an in-chat knowledgebase-MCP step, sincectacommands have no MCP access.Not adopted (deliberately)
No
cas-toolsdependency, no CAS-JSON export, nocap-sc-clientknowledge source, no accession-based hierarchy — evidence in the session showed each to be low value / high cost for this use.Testing
tests/test_negative_markers.py(sign preservation, per-(type,gene)roles incl. the pos-elsewhere/neg-here case, violation thresholding, off-panel handling, dot-plot render, card rendering) andtests/test_cl_id_validation.py(accept/reject + dry-run block/override).🤖 Generated with Claude Code