Skip to content

feat(decision): add optional isolated scoring and reference benchmark - #592

Merged
geisten merged 8 commits into
mainfrom
codex/decision-scoring
Oct 4, 2026
Merged

geisten merged 8 commits into
mainfrom
codex/decision-scoring

Conversation

@geisten

@geisten geisten commented Oct 3, 2026 •

Copy link
Copy Markdown
Owner

Consumers currently have to manage a generation session and raw logits to score alternatives. This adds an isolated numeric decision handle with explicit candidate semantics, reset, ownership and failure behavior.

  • EXPERIMENTAL geist_decision.h; default-off DECISION=1, linkable unsupported stubs, and a content stamp plus forced transitions for reliable on/off builds with coarse timestamps.
  • Per-handle KV/SSM state and creation-sized candidate workspace, immutable shared weights, distinct single-token IDs, model-conformant dense logits and candidate-conditional softmax.
  • Appended architecture vocabulary capability rejects unsupported/embedding-only models before inference. Existing generation, embedding and logits contracts remain intact.
  • Public consumer, hybrid isolation, adversarial numeric/error and real-GGUF tests; CI gates both feature modes in release and sanitizer builds.
  • Reference benchmark records hashes, raw samples, tokenization and hardware/cache conditions, warm p50/p95 and explicitly scoped quality. No trained head or selected-row optimization yet.

Validation on Apple M1 Max:

  • Release and ASan/UBSan unit suites: each 69 passed, 40 skipped, zero failures/errors.
  • Release and sanitizer feature on/off checks; focused Scalar/NEON/Metal tests. Forced transitions also pass with artificially unfavorable output timestamps; unchanged flags keep the build incremental.
  • Real Qwen3 0.6B, Gemma4 E2B softcap and Bonsai 2 27B public scoring parity.
  • Public C23/C++17 headers, API contract smoke, format gate, Python suite.
  • Optimized GCC 16 -O3 -c checks (GCC 15 is not installed) and GCC runtime adversarial tests under fast math. The numeric TUs disable finite-only assumptions; enabled custom builds cannot silently remove NaN/Inf checks. All 23 executed CI checks pass on head 69e91ca (coverage is intentionally skipped for draft PRs).
  • macOS optimized decision disassembly is unchanged by the compiler-flag correction; the recorded smoke baseline remains applicable.

One arithmetic smoke case, CPU NEON / six threads / five prompt tokens / four candidates / FP32 KV / concurrency one / five measured repeats:

Model Decision p50 Dense reference p50 One-token p50 16-token p50
Qwen3 0.6B Q8_0 117.55 ms 119.76 ms 119.30 ms 297.37 ms
Bonsai 2 27B PQ2_0 2974.13 ms 2925.23 ms 2880.10 ms 5641.53 ms

Protocol and limits, all raw samples and hashes. These are workload-specific smoke measurements with uncontrolled caches, not representative task accuracy or Jev parity. Generation quality here measures only its first emitted token; final-answer quality needs downstream evaluation.

Closes #585. Part of #583. Output-row optimization continues in #586; representative Bonsai quality/calibration/latency evaluation in #587.

Expose numeric decisions without making consumers own raw-logit session
state. Each handle shares weights but owns resettable KV/SSM state and
creation-sized candidate storage. Keep default-off, linkable unsupported
stubs; a build stamp and CI gate verify both modes in the same directories.

DENSE preserves model softcaps and defines candidate-conditional
probabilities separately from calibrated confidence. Add public consumer,
adversarial and hybrid isolation tests plus a hashed reference protocol.

Apple M1 Max / cpu_neon / 6 threads, single five-token arithmetic smoke:
Bonsai 2 27B median decision 2.974 s, one-token generation 2.880 s,
16-token generation 5.642 s. This is a reference baseline, not a Jev or
quality-matched speedup claim; selected-row work remains #586.

Refs #583. Closes #585.
@geisten geisten added the decision-inference Optional decision scoring, classifier heads, calibration and quality-matched inference benchmarks label Oct 3, 2026
Linux targets enable -ffast-math, so GCC removed the isfinite guard and
accepted NaN/Inf selected logits. Disable finite-only assumptions for
the numeric decision TU, benchmark reference and adversarial test;
custom enabled builds with finite-only assumptions fail compilation.
Existing inference kernels keep their optimization flags.

GCC 16 runtime tests now reject all three non-finite cases, feature
on/off release and sanitizer tests pass, and optimized macOS decision
code is byte-for-byte the same disassembly as the smoke baseline.
Strict fixture legs reject any unlisted skip. The optional real-model
consumer intentionally skips with DECISION=0, so declare that build
condition. Enabled and disabled public/hybrid tests remain mandatory
through test-decision in release and sanitizer CI.
Keep the benchmark parser's pure size output first, initialize it on
failure, and mark status/size helpers nodiscard per AGENT.md. The parser
family is migrated together; no public signature or inference behavior
changes.
On Intel macOS, decision.o rebuilt with DECISION=0 within the archive's
one-second timestamp window. Make skipped archiving/linking and the
capability witness detected an enabled stale binary. Force those targets
only when the saved flag differs, and keep phony prerequisites out of the
archive's object list. Unchanged-flag builds remain incremental.
@geisten
geisten marked this pull request as ready for review October 4, 2026 10:34
@geisten
geisten merged commit 67fe9f6 into main Oct 4, 2026
24 checks passed
@geisten
geisten deleted the codex/decision-scoring branch October 4, 2026 10:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

decision-inference Optional decision scoring, classifier heads, calibration and quality-matched inference benchmarks

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(decision): optional scoring API and reproducible reference baseline

1 participant