perf(decision): add exact selected-row heads for Apple backends - #599
Merged
Merged
Conversation
Small candidate sets need only a few output rows. Reuse the prompt backbone, normalization and Hadamard basis, then dispatch bounded native row tiles instead of the vocabulary head, sampler and dense logits readback. Keep DENSE as the default reference and DECISION off by default. Reject unsupported layouts explicitly. CPU tiles reuse the same row arithmetic with a restored single-thread OpenMP scope; Metal retains the original n4/n8 pipeline choice. Preallocate tile maps and result buffers and expose actual row work, logical staging bytes and final-stage elapsed time. Release and ASan/UBSan: 70 unit tests passed, 40 expected skips each. Feature transitions, optimized GCC 16, public C/C++ headers, real Qwen, Bonsai/Hadamard and Gemma/softcap parity pass. Before/after generation has byte-identical full logits and 8 token IDs on CPU and Metal. Apple M1 Max / 64 GiB, standalone 4-tile mean head measurements: Qwen-shaped CPU Q8: 3408.4 -> 7.3 us; Bonsai-shaped CPU PQ2: 8302.2 -> 17.5 us; Metal PQ2: 1741.6 -> 350.4 us. Whole-query gains are much smaller: Qwen 1-6%, while Bonsai Metal has no reliable gain in the initial matrix. These are synthetic performance cases, not MMLU quality or Jev parity. The full timing artifact and ordinary prefill/decode sweep follow separately. Refs #586; depends on #585 / #592. MMLU evaluation follows in #587.
cpu_x86 delegates to the Scalar resolver and then replaces its kernels. Importing its row callback incorrectly reported selected-row support and failed on macOS x86. Clear borrowed row slots and bind them only for the owning resolver. A delegating-backend regression fixture reproduces this on every host; x86 public consumers must report unsupported. Release and ASan/UBSan suites: 70 passed, 40 expected skips each; feature transitions, focused Metal/tile sanitizer tests, GCC -O3 and format pass. Apple selected projection arithmetic is unchanged.
…mits Pin raw samples, model/runtime hashes, native row counts, staging bytes, warm allocation counts and process RSS for Qwen and Bonsai on M1 Max / 64 GiB. Preserve ordinary prefill/decode sweeps, reversed-order confirmation and byte-identical full logits/token outputs on Qwen, Bonsai and Gemma. Final standalone Bonsai head: CPU 3359.3 -> 19.4 us (173x), Metal 1939.4 -> 412.7 us (4.7x). Whole-query Bonsai latency remains flat/noisy (0.96-1.01x DENSE/selected); Qwen gains about 1-6%. Report those limits explicitly and defer MMLU quality/calibration to #587.
This was referenced Oct 4, 2026
Comparing the exported resolver address introduced an ARM64 text relocation that Ubuntu's static-to-shared FFI link could not resolve. Use the unique registered backend name to preserve Scalar-only row capability while allowing shared archive consumers to link again. Keep the delegated-backend regression fixture distinct in the registry. Release and ASan/UBSan suites pass (70 passed, 40 expected skips each), as do both feature toggles, format and optimized GCC 16 compilation. Linux ARM64 GCC 14.4 native tests, whole-archive link and Python ctypes load pass with Ubuntu-equivalent PIE settings in an offline container. Apple projection kernels and recorded timings are unchanged.
metal_selected_rows only ensured the Q4_K pipelines before encoding. An encoder whose pipeline is missing dispatches nothing, so the query returned OK with the previous query's tile values in y. Check the dtype's pipeline with metal_selected_prepare first and return GEIST_E_UNSUPPORTED instead. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AD86DiU1PpDqYprALdARAw
geisten
marked this pull request as ready for review
October 4, 2026 10:36
geisten
enabled auto-merge
October 4, 2026 10:36
This was referenced Oct 4, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Small candidate sets still pay for the full vocabulary projection in the decision API. Add explicit
GEIST_DECISION_SELECTED_ROWSmode to reuse the prompt backbone and final normalization/rotation, then project bounded native row tiles. Candidate logits and conditional probabilities retain the DENSE contract; unsupported model/backend pairs fail explicitly.CPU Scalar/Apple NEON source rows and Bonsai PQ2 x8 reuse resolved kernels. Tiny CPU tiles avoid repeated OpenMP team starts and restore the calling task's thread setting. Scalar resolver delegation cannot export its row capability to a backend that replaces those kernels. Metal retains the original dense n4/n8 arithmetic, including tails. Readouts preallocate their workspace and report projected rows, logical staging bytes and final-stage elapsed time. DENSE remains the default and
DECISION=0remains the default build.Validation:
Measurement artifact and raw samples pin 400 decision samples, model/runtime hashes, row work, staging bytes, zero warm selected-kernel heap allocations and process RSS. On M1 Max / 64 GiB, final standalone four-tile Bonsai-shaped PQ2 means are 3359.3 → 19.4 µs on CPU (173x) and 1939.4 → 412.7 µs on Metal (4.7x).
Whole-query gains are much smaller: Qwen gains about 1–6%; Bonsai remains flat/noisy and sometimes a few percent slower. The backbone dominates. Synthetic timings establish no MMLU quality, calibration, quality-matched reasoning speedup or Jev parity. Private sessions still reserve ordinary dense scratch, and GPU final-stage time can include pending backbone work.
Depends on #592 (#585): this PR is stacked on
codex/decision-scoringto isolate #586. Retarget tomainonce the base PR lands. MMLU-first evaluation on Apple Silicon is recorded in #587; other backends and trained classifier heads remain later work.Closes #586. Refs #583.