You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Parent: #583
Depends on: #585
Use #586 for the final optimized comparison; collect the reference results before that optimization.
Planned branch: codex/decision-evaluation.
Problem / motivation
Removing generated tokens can yield a large speedup, but Bonsai's published thinking-mode scores do not establish its immediate-decision quality. We need measured results before selecting classifier training or making Jev-like speedup claims.
Scope
Confirmed rollout: start on Apple Silicon (cpu_neon and Metal), with MMLU first. Other devices/backends follow later. Pin the MMLU version/split and prompt/answer-token construction, and declare the generative quality baseline and non-inferiority margin before any quality-matched speedup comparison. The synthetic #586 timing matrix is performance-only and supplies no MMLU accuracy evidence.
Build on the current Bonsai implementation, existing eval_geist/MMLU adapters and the protocol established in #585. Keep reusable runtime evaluation mechanisms here; application policy and domain datasets belong to the consumer.
Compare on identical held-out questions:
Direct answer-token scores without generated reasoning.
An optimized single-answer-token generative baseline.
Ordinary generative/reasoning decisions, counting every generated reasoning and answer token.
A hosted Jev run is optional only when account access and benchmark spending are explicitly authorized. Published provider speedups are contextual references, not local measurements or acceptance thresholds.
Use separate development, calibration and held-out test data with documented split policy. Include representative negative/ambiguous cases and label-order sensitivity.
Report accuracy or macro-F1, log-loss/Brier score and reliability/calibration results, plus p50/p95 latency, throughput, memory and cold/warm behavior with raw samples.
Any post-hoc temperature scaling is fitted only on the calibration split and tied to the checkpoint, precision/backend policy and prompt construction. Uncalibrated option probabilities remain clearly identified.
Define the quality target/non-inferiority margin before comparing speedups. Report speedup at that target; report any quality loss instead of trading it away silently.
Count preprocessing, all model calls and any abstention/fallback route in the reported end-to-end policy.
Conclude whether token scoring is adequate, a trained head is justified, or backbone adaptation is needed. Record negative findings too.
Existing dependencies / history
PQ2_0 and Hadamard support are already implemented. #432 and #561 are closed fixes. Link any newly reproduced backend issue to its existing owner rather than duplicating Bonsai support. Reuse #364 for performance methodology.
Parent: #583
Depends on: #585
Use #586 for the final optimized comparison; collect the reference results before that optimization.
Planned branch:
codex/decision-evaluation.Problem / motivation
Removing generated tokens can yield a large speedup, but Bonsai's published thinking-mode scores do not establish its immediate-decision quality. We need measured results before selecting classifier training or making Jev-like speedup claims.
Scope
Confirmed rollout: start on Apple Silicon (
cpu_neonand Metal), with MMLU first. Other devices/backends follow later. Pin the MMLU version/split and prompt/answer-token construction, and declare the generative quality baseline and non-inferiority margin before any quality-matched speedup comparison. The synthetic #586 timing matrix is performance-only and supplies no MMLU accuracy evidence.Build on the current Bonsai implementation, existing
eval_geist/MMLU adapters and the protocol established in #585. Keep reusable runtime evaluation mechanisms here; application policy and domain datasets belong to the consumer.Compare on identical held-out questions:
A hosted Jev run is optional only when account access and benchmark spending are explicitly authorized. Published provider speedups are contextual references, not local measurements or acceptance thresholds.
Acceptance
Existing dependencies / history
PQ2_0 and Hadamard support are already implemented. #432 and #561 are closed fixes. Link any newly reproduced backend issue to its existing owner rather than duplicating Bonsai support. Reuse #364 for performance methodology.