Goal
Add optional, application-neutral decision inference to geistlib: score explicit alternatives or trained classifier outputs without generating an answer string. Measure speedups at a declared, matched decision-quality target; do not treat Jev's advertised workflow ratios as a universal acceptance threshold.
Current foundation
As of main d2641c9, Bonsai 2 support already exists: PQ2_0, prism.hadamard, backend kernels and end-to-end coverage. Reuse it rather than creating a second format implementation.
Architecture and rollout
- An additive, EXPERIMENTAL decision API with its own handle/options; immutable model weights remain shared and mutable state/workspace is owned by each decision instance.
- A new decision build feature, disabled by default initially (proposed build spelling:
DECISION=1), plus explicit instance-level readout selection. Existing Bonsai support is already shipped and must not become disabled by default as a side effect.
- Start with selected answer-token scoring. Add an optimized selected-row readout behind the same contract after reference parity is established.
- Resolve capabilities and allocate workspace at creation/plan time. Unsupported modes return
GEIST_E_UNSUPPORTED; no silent fallback that changes score semantics.
- Prompts, chat templates, JSON schemas, routing policy and application datasets remain with consumers, per
docs/README.md. The library exposes numeric inference mechanisms.
- Trained heads and multi-question evaluation are separate, gated research items.
Delivery
Use one bounded branch/PR per implementable issue, based on current main. The initial branch is codex/decision-scoring. Merge complete pieces behind the disabled decision feature; avoid a long-lived integration branch. Organize with the decision-inference and bonsai labels and this tracking issue; no separate GitHub Project.
Implementation
Delivered: #585 was merged in #592, and #586 was merged in #599. The additive API and selected-row readout remain behind the default-off DECISION=1 build feature. Evaluation infrastructure for #587 is in draft PR #608 on codex/decision-evaluation. The approved eight-question classic-MMLU development pilot is complete in #608: Metal first, then Apple CPU NEON, two trials per arm, cap512 and 128 calls/backend. Exact selected/DENSE logits and deterministic repeats pass. Direct chat and reasoning each get 4/8; cloze gets 6/8; reasoning has four invalid capped answers per backend. Descriptive timing ratios are 2.168 (Metal) and 2.328 (CPU). Held-out quality-matched speedup and Jev equivalence remain unproven; #587 and #588 stay open for their remaining evidence.
Confirmed evaluation order: Apple Silicon first (cpu_neon and Metal), MMLU first in #587; other backends follow later. The #586 measurement artifact documents strong head-only savings but no reliable Bonsai whole-query improvement. Quality-matched speedup remains unproven until #587.
Deferred research
Existing Bonsai work
These existing issues are related history, not duplicate implementation tasks or newly reopened blockers. Revalidate the relevant behavior as part of #587.
Definition of done
- Feature-on and feature-off configurations have explicit, tested API behavior.
- Existing text generation, logits access and embedding inference retain their contracts and quality.
- Score semantics distinguish full-vocabulary probabilities, probabilities conditional on supplied alternatives, and empirically calibrated correctness estimates.
- Benchmarks pin model/runtime revisions, tokenizer/prompt construction, hardware/backend, input/output lengths, question count, cache state and concurrency. Publish raw results, p50/p95, memory, decision quality and calibration where applicable.
- Compare against both an optimized one-token/score baseline and ordinary generative/reasoning workflows. Report negative results and the cost of any fallback.
- Follow AGENT.md: checked geometry, defined failure outputs, no hot-path allocation, required release/sanitizer/gcc checks and performance gates for each relevant implementation PR.
References
Goal
Add optional, application-neutral decision inference to geistlib: score explicit alternatives or trained classifier outputs without generating an answer string. Measure speedups at a declared, matched decision-quality target; do not treat Jev's advertised workflow ratios as a universal acceptance threshold.
Current foundation
As of main
d2641c9, Bonsai 2 support already exists:PQ2_0,prism.hadamard, backend kernels and end-to-end coverage. Reuse it rather than creating a second format implementation.tools/eval_geist.c(SCOREALT) andgeist_session_peek_logits.Architecture and rollout
DECISION=1), plus explicit instance-level readout selection. Existing Bonsai support is already shipped and must not become disabled by default as a side effect.GEIST_E_UNSUPPORTED; no silent fallback that changes score semantics.docs/README.md. The library exposes numeric inference mechanisms.Delivery
Use one bounded branch/PR per implementable issue, based on current main. The initial branch is
codex/decision-scoring. Merge complete pieces behind the disabled decision feature; avoid a long-lived integration branch. Organize with thedecision-inferenceandbonsailabels and this tracking issue; no separate GitHub Project.Implementation
Delivered: #585 was merged in #592, and #586 was merged in #599. The additive API and selected-row readout remain behind the default-off
DECISION=1build feature. Evaluation infrastructure for #587 is in draft PR #608 oncodex/decision-evaluation. The approved eight-question classic-MMLU development pilot is complete in #608: Metal first, then Apple CPU NEON, two trials per arm, cap512 and 128 calls/backend. Exact selected/DENSE logits and deterministic repeats pass. Direct chat and reasoning each get 4/8; cloze gets 6/8; reasoning has four invalid capped answers per backend. Descriptive timing ratios are 2.168 (Metal) and 2.328 (CPU). Held-out quality-matched speedup and Jev equivalence remain unproven; #587 and #588 stay open for their remaining evidence.Confirmed evaluation order: Apple Silicon first (
cpu_neonand Metal), MMLU first in #587; other backends follow later. The #586 measurement artifact documents strong head-only savings but no reliable Bonsai whole-query improvement. Quality-matched speedup remains unproven until #587.Deferred research
Existing Bonsai work
3f86e33,406abc4).These existing issues are related history, not duplicate implementation tasks or newly reopened blockers. Revalidate the relevant behavior as part of #587.
Definition of done
References