Skip to content

eval(decision): Bonsai 2 quality, calibration and latency without generated reasoning #587

Description

@geisten

Parent: #583
Depends on: #585
Use #586 for the final optimized comparison; collect the reference results before that optimization.
Planned branch: codex/decision-evaluation.

Problem / motivation

Removing generated tokens can yield a large speedup, but Bonsai's published thinking-mode scores do not establish its immediate-decision quality. We need measured results before selecting classifier training or making Jev-like speedup claims.

Scope

Confirmed rollout: start on Apple Silicon (cpu_neon and Metal), with MMLU first. Other devices/backends follow later. Pin the MMLU version/split and prompt/answer-token construction, and declare the generative quality baseline and non-inferiority margin before any quality-matched speedup comparison. The synthetic #586 timing matrix is performance-only and supplies no MMLU accuracy evidence.

Build on the current Bonsai implementation, existing eval_geist/MMLU adapters and the protocol established in #585. Keep reusable runtime evaluation mechanisms here; application policy and domain datasets belong to the consumer.

Compare on identical held-out questions:

  1. Direct answer-token scores without generated reasoning.
  2. An optimized single-answer-token generative baseline.
  3. Ordinary generative/reasoning decisions, counting every generated reasoning and answer token.
  4. The selected-row optimization once perf(decision): compute only requested output rows behind the scoring API #586 is ready.

A hosted Jev run is optional only when account access and benchmark spending are explicitly authorized. Published provider speedups are contextual references, not local measurements or acceptance thresholds.

Acceptance

  • Pin Bonsai checkpoint/hash, runtime revisions, tokenizer/template, backend/device, numerical modes, candidate ordering, input lengths, question count, cache policy and concurrency.
  • Recheck same-token runtime parity against a pinned Prism reference using the existing coverage from Bonsai-27B: MMLU 0.44 against the PrismML fork's 0.76 on identical tokens #432; include longer/high-entropy cases, not only a short greedy smoke test.
  • Use separate development, calibration and held-out test data with documented split policy. Include representative negative/ambiguous cases and label-order sensitivity.
  • Report accuracy or macro-F1, log-loss/Brier score and reliability/calibration results, plus p50/p95 latency, throughput, memory and cold/warm behavior with raw samples.
  • Any post-hoc temperature scaling is fitted only on the calibration split and tied to the checkpoint, precision/backend policy and prompt construction. Uncalibrated option probabilities remain clearly identified.
  • Define the quality target/non-inferiority margin before comparing speedups. Report speedup at that target; report any quality loss instead of trading it away silently.
  • Count preprocessing, all model calls and any abstention/fallback route in the reported end-to-end policy.
  • Conclude whether token scoring is adequate, a trained head is justified, or backbone adaptation is needed. Record negative findings too.

Existing dependencies / history

PQ2_0 and Hadamard support are already implemented. #432 and #561 are closed fixes. Link any newly reproduced backend issue to its existing owner rather than duplicating Bonsai support. Reuse #364 for performance methodology.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bonsaiBonsai model support, Prism ternary formats, correctness and performancedecision-inferenceOptional decision scoring, classifier heads, calibration and quality-matched inference benchmarksenhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions