Skip to content

Tracking: modular decision inference in geistlib #583

Description

@geisten

Goal

Add optional, application-neutral decision inference to geistlib: score explicit alternatives or trained classifier outputs without generating an answer string. Measure speedups at a declared, matched decision-quality target; do not treat Jev's advertised workflow ratios as a universal acceptance threshold.

Current foundation

As of main d2641c9, Bonsai 2 support already exists: PQ2_0, prism.hadamard, backend kernels and end-to-end coverage. Reuse it rather than creating a second format implementation.

Architecture and rollout

  • An additive, EXPERIMENTAL decision API with its own handle/options; immutable model weights remain shared and mutable state/workspace is owned by each decision instance.
  • A new decision build feature, disabled by default initially (proposed build spelling: DECISION=1), plus explicit instance-level readout selection. Existing Bonsai support is already shipped and must not become disabled by default as a side effect.
  • Start with selected answer-token scoring. Add an optimized selected-row readout behind the same contract after reference parity is established.
  • Resolve capabilities and allocate workspace at creation/plan time. Unsupported modes return GEIST_E_UNSUPPORTED; no silent fallback that changes score semantics.
  • Prompts, chat templates, JSON schemas, routing policy and application datasets remain with consumers, per docs/README.md. The library exposes numeric inference mechanisms.
  • Trained heads and multi-question evaluation are separate, gated research items.

Delivery

Use one bounded branch/PR per implementable issue, based on current main. The initial branch is codex/decision-scoring. Merge complete pieces behind the disabled decision feature; avoid a long-lived integration branch. Organize with the decision-inference and bonsai labels and this tracking issue; no separate GitHub Project.

Implementation

Delivered: #585 was merged in #592, and #586 was merged in #599. The additive API and selected-row readout remain behind the default-off DECISION=1 build feature. Evaluation infrastructure for #587 is in draft PR #608 on codex/decision-evaluation. The approved eight-question classic-MMLU development pilot is complete in #608: Metal first, then Apple CPU NEON, two trials per arm, cap512 and 128 calls/backend. Exact selected/DENSE logits and deterministic repeats pass. Direct chat and reasoning each get 4/8; cloze gets 6/8; reasoning has four invalid capped answers per backend. Descriptive timing ratios are 2.168 (Metal) and 2.328 (CPU). Held-out quality-matched speedup and Jev equivalence remain unproven; #587 and #588 stay open for their remaining evidence.

Confirmed evaluation order: Apple Silicon first (cpu_neon and Metal), MMLU first in #587; other backends follow later. The #586 measurement artifact documents strong head-only savings but no reliable Bonsai whole-query improvement. Quality-matched speedup remains unproven until #587.

Deferred research

Existing Bonsai work

These existing issues are related history, not duplicate implementation tasks or newly reopened blockers. Revalidate the relevant behavior as part of #587.

Definition of done

  • Feature-on and feature-off configurations have explicit, tested API behavior.
  • Existing text generation, logits access and embedding inference retain their contracts and quality.
  • Score semantics distinguish full-vocabulary probabilities, probabilities conditional on supplied alternatives, and empirically calibrated correctness estimates.
  • Benchmarks pin model/runtime revisions, tokenizer/prompt construction, hardware/backend, input/output lengths, question count, cache state and concurrency. Publish raw results, p50/p95, memory, decision quality and calibration where applicable.
  • Compare against both an optimized one-token/score baseline and ordinary generative/reasoning workflows. Report negative results and the cost of any fallback.
  • Follow AGENT.md: checked geometry, defined failure outputs, no hot-path allocation, required release/sanitizer/gcc checks and performance gates for each relevant implementation PR.

References

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bonsaiBonsai model support, Prism ternary formats, correctness and performancedecision-inferenceOptional decision scoring, classifier heads, calibration and quality-matched inference benchmarksenhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions