Skip to content

perf(decision): add exact selected-row heads for Apple backends - #599

Merged
geisten merged 6 commits into
mainfrom
codex/decision-fast-path
Oct 4, 2026
Merged

geisten merged 6 commits into
mainfrom
codex/decision-fast-path

Conversation

@geisten

@geisten geisten commented Oct 4, 2026 •

Copy link
Copy Markdown
Owner

Small candidate sets still pay for the full vocabulary projection in the decision API. Add explicit GEIST_DECISION_SELECTED_ROWS mode to reuse the prompt backbone and final normalization/rotation, then project bounded native row tiles. Candidate logits and conditional probabilities retain the DENSE contract; unsupported model/backend pairs fail explicitly.

CPU Scalar/Apple NEON source rows and Bonsai PQ2 x8 reuse resolved kernels. Tiny CPU tiles avoid repeated OpenMP team starts and restore the calling task's thread setting. Scalar resolver delegation cannot export its row capability to a backend that replaces those kernels. Metal retains the original dense n4/n8 arithmetic, including tails. Readouts preallocate their workspace and report projected rows, logical staging bytes and final-stage elapsed time. DENSE remains the default and DECISION=0 remains the default build.

Validation:

  • Release and ASan/UBSan unit suites: 70 passed, 40 expected skips each; both feature transitions pass.
  • Linux ARM64 GCC 14.4.0 native build with Ubuntu-equivalent PIE settings: decision/error/tile fixtures, whole-archive shared-library link and Python ctypes load pass.
  • Optimized GCC 16 builds and non-finite/error runtime tests; format, C23/C++17 public headers and the 18-symbol stable API contract pass.
  • Bit-identical dense/selected native tiles and real Qwen, Bonsai PQ2/Hadamard and Gemma softcap checks on Apple CPU/Metal, including sanitizer coverage and a portable resolver-delegation regression fixture.
  • Before/after ordinary generation has byte-identical full logits and eight generated token IDs on all three models and both Apple backends. Standard 32/128-token prefill/decode sweeps and reversed-order confirmation are retained.

Measurement artifact and raw samples pin 400 decision samples, model/runtime hashes, row work, staging bytes, zero warm selected-kernel heap allocations and process RSS. On M1 Max / 64 GiB, final standalone four-tile Bonsai-shaped PQ2 means are 3359.3 → 19.4 µs on CPU (173x) and 1939.4 → 412.7 µs on Metal (4.7x).

Whole-query gains are much smaller: Qwen gains about 1–6%; Bonsai remains flat/noisy and sometimes a few percent slower. The backbone dominates. Synthetic timings establish no MMLU quality, calibration, quality-matched reasoning speedup or Jev parity. Private sessions still reserve ordinary dense scratch, and GPU final-stage time can include pending backbone work.

Depends on #592 (#585): this PR is stacked on codex/decision-scoring to isolate #586. Retarget to main once the base PR lands. MMLU-first evaluation on Apple Silicon is recorded in #587; other backends and trained classifier heads remain later work.

Closes #586. Refs #583.

Small candidate sets need only a few output rows. Reuse the prompt backbone,
normalization and Hadamard basis, then dispatch bounded native row tiles
instead of the vocabulary head, sampler and dense logits readback.

Keep DENSE as the default reference and DECISION off by default. Reject
unsupported layouts explicitly. CPU tiles reuse the same row arithmetic
with a restored single-thread OpenMP scope; Metal retains the original
n4/n8 pipeline choice. Preallocate tile maps and result buffers and expose
actual row work, logical staging bytes and final-stage elapsed time.

Release and ASan/UBSan: 70 unit tests passed, 40 expected skips each.
Feature transitions, optimized GCC 16, public C/C++ headers, real Qwen,
Bonsai/Hadamard and Gemma/softcap parity pass. Before/after generation has
byte-identical full logits and 8 token IDs on CPU and Metal.

Apple M1 Max / 64 GiB, standalone 4-tile mean head measurements:
Qwen-shaped CPU Q8: 3408.4 -> 7.3 us; Bonsai-shaped CPU PQ2: 8302.2 ->
17.5 us; Metal PQ2: 1741.6 -> 350.4 us. Whole-query gains are much smaller:
Qwen 1-6%, while Bonsai Metal has no reliable gain in the initial matrix.
These are synthetic performance cases, not MMLU quality or Jev parity.
The full timing artifact and ordinary prefill/decode sweep follow separately.

Refs #586; depends on #585 / #592. MMLU evaluation follows in #587.
cpu_x86 delegates to the Scalar resolver and then replaces its kernels. Importing its row callback incorrectly reported selected-row support and failed on macOS x86. Clear borrowed row slots and bind them only for the owning resolver. A delegating-backend regression fixture reproduces this on every host; x86 public consumers must report unsupported.

Release and ASan/UBSan suites: 70 passed, 40 expected skips each; feature transitions, focused Metal/tile sanitizer tests, GCC -O3 and format pass. Apple selected projection arithmetic is unchanged.
…mits

Pin raw samples, model/runtime hashes, native row counts, staging bytes, warm allocation counts and process RSS for Qwen and Bonsai on M1 Max / 64 GiB. Preserve ordinary prefill/decode sweeps, reversed-order confirmation and byte-identical full logits/token outputs on Qwen, Bonsai and Gemma.

Final standalone Bonsai head: CPU 3359.3 -> 19.4 us (173x), Metal 1939.4 -> 412.7 us (4.7x). Whole-query Bonsai latency remains flat/noisy (0.96-1.01x DENSE/selected); Qwen gains about 1-6%. Report those limits explicitly and defer MMLU quality/calibration to #587.
@geisten geisten added decision-inference Optional decision scoring, classifier heads, calibration and quality-matched inference benchmarks bonsai Bonsai model support, Prism ternary formats, correctness and performance labels Oct 4, 2026
geisten and others added 2 commits October 4, 2026 10:20
Comparing the exported resolver address introduced an ARM64 text
relocation that Ubuntu's static-to-shared FFI link could not resolve.
Use the unique registered backend name to preserve Scalar-only row
capability while allowing shared archive consumers to link again.

Keep the delegated-backend regression fixture distinct in the registry.
Release and ASan/UBSan suites pass (70 passed, 40 expected skips each),
as do both feature toggles, format and optimized GCC 16 compilation.
Linux ARM64 GCC 14.4 native tests, whole-archive link and Python ctypes
load pass with Ubuntu-equivalent PIE settings in an offline container.
Apple projection kernels and recorded timings are unchanged.
metal_selected_rows only ensured the Q4_K pipelines before encoding. An
encoder whose pipeline is missing dispatches nothing, so the query
returned OK with the previous query's tile values in y. Check the
dtype's pipeline with metal_selected_prepare first and return
GEIST_E_UNSUPPORTED instead.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AD86DiU1PpDqYprALdARAw
Base automatically changed from codex/decision-scoring to main October 4, 2026 10:34
@geisten
geisten marked this pull request as ready for review October 4, 2026 10:36
@geisten
geisten enabled auto-merge October 4, 2026 10:36
@geisten
geisten merged commit 5ec0b0f into main Oct 4, 2026
24 checks passed
@geisten
geisten deleted the codex/decision-fast-path branch October 4, 2026 10:58
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bonsai Bonsai model support, Prism ternary formats, correctness and performance decision-inference Optional decision scoring, classifier heads, calibration and quality-matched inference benchmarks

Projects

None yet

Development

Successfully merging this pull request may close these issues.

perf(decision): compute only requested output rows behind the scoring API

2 participants