Skip to content

feat: device-agnostic execution + GGUF Q8_0 quantized serving + CPU decode/prefill perf - #1917

Merged
ooples merged 19 commits into
feat/serving-integrationfrom
feat/device-agnostic-model
Jul 21, 2026
Merged

ooples merged 19 commits into
feat/serving-integrationfrom
feat/device-agnostic-model

Conversation

@ooples

@ooples ooples commented Jul 21, 2026

Copy link
Copy Markdown
Owner

Summary

Device-agnostic execution + GGUF Q8_0 quantized serving + CPU decode/prefill performance, stacked on the unified continuous-batching engine (#1888).

Verified end-to-end on real SmolLM2-135M GGUF, llama.cpp parity preserved throughout (greedy next-token = 7042 " Paris" at every step).

Performance (SmolLM2-135M Q8_0, this AVX2 box)

before after
Prefill 205 tok/s 451 (2.2×)
Decode 3.9 tok/s 11.3 (2.9×)

What's here

  • Q8_0 quantized serving: keep GGUF Q8_0 weights quantized instead of expanding to fp32 at load (GgufFile.TryReadQ8_0Raw -> GgufModelSource.TryReadQ8_0 -> LlamaModelBuilder.LoadDense -> DenseLayer.SetQuantizedWeightsQ8_0), running the block-Q8_0 GEMM (Tensors Q8BlockGemm) directly on int8. 3.76x less weight RAM. The quantized forward is gated on VNNI hardware (AvxVnni/Avx512) — on this AVX2 box it stays on fp32 (the int8 dot needs vpdpbusd to beat the tuned fp32 BLAS), auto-engaging on the CUDA+Linux serving target.
  • Decode perf: vectorized the scalar per-token projection matvec (PagedCachedMultiHeadAttention.MatVecMul, AVX2 FMA); bulk-read the paged KV history under one lock (PagedKVCache.ReadKeyValueRange) instead of a lock per position.
  • Device-agnostic API (from the branch): model.To(DeviceInfo) / layer.To(...) placement, ConfigureModel(source, DeviceInfo) facade overload, GQA paged incremental wiring (InferenceOptimizer recurses into PreLNTransformerBlock).
  • Benchmark harnesses: DEVHOST_PROFILE (steady-state prefill for dotnet-trace) and DEVHOST_DECODE (real incremental generation -> tok/s).

Release gate

The pin is a local probe (AiDotNet.Tensors 0.118.10-q8) — this consumer work depends on Tensors feat/q8-block-gemm (the Q8BlockGemm kernel + the float-RoPE/NumOps/mask-cache perf fixes, opened as a separate PR) plus #828 (already merged). CI will be red until that Tensors branch is released to nuget.org and the pin is bumped to the released version.

🤖 Generated with Claude Code

ooples and others added 19 commits July 20, 2026 09:57
PyTorch-style device placement, enum-only (no magic-string overloads). LayerBase.To
moves registered parameters + buffers and recurses sublayers; NeuralNetworkBase.To
moves every layer. Skips zero-length (deferred) params. First user-facing piece of
the device-agnostic execution model; builds net10.0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…el 1)

Load a pretrained checkpoint and place the whole model on a device in one call —
the common "load onto my GPU" case — via the type-safe enum DeviceInfo (no device
strings). Additive overload: existing ConfigureModel(source) callers are unchanged.
Delegates to model.To(device). Builds net10.0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Proves the PyTorch-style placement API (model.To(DeviceInfo.OpenCL())) drives
correct GPU execution to llama.cpp's greedy token. Uses AutoDetectAndConfigureGpu
so placement (Tensor.To -> global backend) and execution (Current dispatcher)
share one backend; skips when no GPU is wired in-process (validates on DevHost/CUDA).
Surfaced that To() and Current use different backend acquisition — unify in Phase 1.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
… gap

Runs the whole decoder forward inside a DeferredScope (BeginDeferredScope ->
Predict -> Execute) as the acceptance check for transparent fusion. Today it
returns token 0 (all-zero logits) instead of 7042: the graph capture/replay does
not materialize the decoder's output, so transparent fusion needs the graph path
(output binding + RoPE/GQA-SDPA/RMSNorm recording) debugged before it can wrap
Predict. Skipped with that reason; un-skip when the graph path is fixed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
… for GQA KV-cache

Adds PreLNTransformerBlock.ReplaceAttention (mirrors TransformerEncoderBlock) and a
PreLNTransformerBlock case in ApplyAttentionOptimizations that swaps the block's nested
GroupedQueryAttentionLayer for a KV-cached CachedGroupedQueryAttention (shared helper
BuildCachedGqaReplacement, reused by the top-level GQA case). This is the first piece of
wiring the GGUF/LLaMA decoder (GQA nested in PreLNTransformerBlock) to incremental KV-cache
decode. Still inert end-to-end: InitializeGQAKVCache's collection scan is also top-level
(won't yet find the nested cached GQA), and ServableModelWrapper requires a paged cache
that GQA lacks — so the incremental clone is still discarded and serving falls back to the
eager model (no regression). Remaining: recurse InitializeGQAKVCache; accept the non-paged
GQA cache as the incremental cache. Builds net10.0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…attention

The shared attention-host enumerator (used by the GQA/paged KV-cache init,
attention-optimizability detection, and quantization scans) recursed into
TransformerEncoderBlock/TransformerDecoderBlock but not PreLNTransformerBlock,
where LLaMA/GGUF decoders host their grouped-query attention. Add that case so
InitializeGQAKVCache (and the other scans) find the nested CachedGroupedQueryAttention.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add PagedAttentionConfig.NumQueryHeads (0 => same as NumHeads = MHA). NumHeads is
now the KV/cache head count; query heads may outnumber it under GQA, with each KV
head shared by NumQueryHeads/NumHeads query heads (kvHead = qHead / group). Threaded
through every compute path — ComputeAttention, ComputeTiledPagedAttention (decode),
ComputeContiguousCausalPrefill (prefill), ComputeBatchedAttention, and the fused
Forward/ForwardQuantized (asymmetric q_proj vs k_proj/v_proj, matching HF). K/V
buffers size to the KV heads the cache stores; per-query-head online-softmax scratch.
MHA stays the group==1 special case (all 43 PagedAttention tests unchanged).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…tion aware

Add kvHeadCount ctor param (0 => headCount = MHA). K/V weights are now
[embDim, kvHeadCount*headDim] (narrower under GQA); Q/O stay [embDim, embDim].
Threaded kvProjDim through every projection + cache write in the per-token,
batched-GEMM, and contiguous-prefill paths; RoPE applies to Q over query heads
and K over KV heads; param serialization and int8 quant use the real per-weight
widths. Expose KVHeadCount for the optimizer to build the paged config/kernel.

Head split/merge and repeat_kv now use vectorized library ops (Reshape/Transpose
and Engine.TensorGather on the head axis) instead of scalar nested loops. MHA is
the kvHeadCount==headCount special case; all 43 PagedAttention tests unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
InitializePagedKVCache now sizes the PagedKVCache by the KV-head count and passes
NumQueryHeads to the kernel so it repeats each KV head across its query-head group.
BuildPagedGqaReplacement converts a GroupedQueryAttentionLayer (top-level or nested
in PreLNTransformerBlock) to a PagedCachedMultiHeadAttention with the source KV-head
count, copying its [Q][K][V][O][outBias] parameters and RoPE/ALiBi config; the GQA
branches prefer it when paged KV is enabled. Guarded by GroupedQueryAttentionLayer.
UsesProjectionBias: models with Q/K/V projection bias (Qwen2-style) fall back to the
contiguous CachedGroupedQueryAttention, which the paged layer cannot represent.

200 paged/optimizer/KVCache/GQA tests green net10.0; both TFMs build.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…oder

Numerical-parity test: a tiny GQA decoder (numHeads=4, numKVHeads=2, RoPE) nested
in a PreLNTransformerBlock is optimized with the paged KV cache enabled, and the
optimized forward matches the original within 1e-3 — proving the paged-GQA kernel
repeat-KV, the layer's narrow K/V projections, the [Q][K][V][O][outBias] weight
copy, and the interleaved-RoPE convention are all faithful. Also asserts the GQA is
rewritten to PagedCachedMultiHeadAttention and the paged KV cache is live.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…as, softcap

A cloned GQA layer (serialize -> deserialize) silently lost its RoPE positional
encoding, causal mask, custom head dimension, Q/K/V projection bias, and attention
logit soft-cap: GetMetadata never persisted them and the deserializer defaulted the
bools to false with no ConfigurePositionalEncoding call. So a cloned decoder computed
bidirectional, RoPE-less attention and diverged — which is exactly what breaks the
incremental-serving clone of a GGUF/LLaMA decoder. Persist all of them in GetMetadata
and restore them (incl. RoPE via the layer's ConfigurePositionalEncoding) on deserialize.

New test clones a tiny RoPE+causal GQA model and asserts forward parity within 1e-4.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…g engages

The block hosts a polymorphic (T5/MHA/GQA) attention sublayer, so it had no
deserialization constructor and cloning threw — which silently disabled the paged
incremental-generation clone (ServableModelWrapper.BuildIncrementalModel) for every
GGUF/LLaMA decoder. GetMetadata now persists the block dims + FFN activation + the
nested attention as an 'Attn.'-prefixed self-contained sub-blob (type + its metadata
+ shapes); a new DeserializationHelper branch rebuilds the attention recursively via
CreateLayerFromType (E1 keeps its RoPE/mask) and constructs the block. OnFirstForward
resolves the lazy norms + FFN in forward order (deser/resolve path only, so no RNG
perturbation on a normal forward) so ParameterCount is correct before SetParameters.

The paged-GQA parity test now runs with cloneModel: true (the real serving path) and
still matches the original within 1e-3.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
End-to-end proof that paged-GQA engages in serving: a numHeads>numKVHeads PreLN
decoder (RoPE) wrapped in ServableModelWrapper now reports SupportsIncrementalGeneration
true and generates in-range tokens. Exercises the full chain — GQA-aware paged kernel,
narrow-K/V paged attention layer, optimizer GQA->paged rewrite, and PreLNTransformerBlock
+ GQA serialization (the clone BuildIncrementalModel performs). Before this the clone
threw and every LLaMA/GGUF-style decoder silently fell back to the stateless path.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…engine ops [prototype]

Rewrites the inference-only cached GQA layer to apply RoPE via Engine.ApplyRoPEInterleaved
and grouped-query attention via Engine.ScaledDotProductAttentionGqa (unexpanded K/V) — a
device-agnostic, GPU-graph-recordable forward that drops the two managed unrecordable
hotspots (managed RoPE + ExpandKVHeads). RotaryPositionalEncodingLayer exposes its cos/sin
caches (GetInterleavedCaches) for the fused op. ALiBi keeps the expand + FlashAttention path.

Verified: the optimized (non-paged) model's forward matches the original managed GQA decoder
within 1e-3 (InferenceOptimizer_CachedGroupedQueryAttention_RewrittenForward_MatchesOriginal),
run against a local AiDotNet.Tensors probe.

PROTOTYPE / RELEASE-GATED: the pin is a LOCAL probe (0.117.200-devexec) of the unreleased
AiDotNet.Tensors PR #828 (device-agnostic execution). It MUST be bumped to the released
AiDotNet.Tensors version before this lands / any AiDotNet PR — CI cannot restore the probe.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…offset fix)

Local probe of the Tensors branch after the OpenCL GQA-SDPA causal offset fix. Still
release-gated — bumps to the released AiDotNet.Tensors before any AiDotNet PR.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
… + profiling harness

keeps GGUF Q8_0 linear weights in their native int8 layout instead of expanding to
fp32 at load, and runs the block-Q8_0 GEMM (Tensors Q8BlockGemm, ggml
q8_0 parity) directly on them for float inference. GgufFile.TryReadQ8_0Raw reads
the native blocks (int8 + per-32 fp16 scale); GgufModelSource.TryReadQ8_0 maps the
HF name; LlamaModelBuilder.LoadDense installs them on the FFN + lm_head DenseLayers;
DenseLayer runs the quantized GEMM on the float inference path. gated to small M
(decode/small-batch, where the naive int8 kernel is bandwidth-bound and beats fp32
BLAS); large-M prefill stays on fp32 until the int8 GEMM is register-tiled, so
prefill is unchanged (451 tok/s) while decode runs on int8. gguf 7042 llama.cpp
parity preserved. also adds the DEVHOST_PROFILE steady-state loop used to profile
the forward. NOTE: release-gated on Tensors Q8BlockGemm (probe pin 0.118.10-q8).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
the block-Q8_0 GEMM only beats the tuned fp32 BLAS when the CPU has a 1-instruction
VNNI int8 multiply-accumulate (AvxVnni / Avx512 vpdpbusd). on AVX2-only CPUs the
3-instruction maddubs dot is no faster, so engaging it there would regress. gate the
DenseLayer quantized path on AvxVnni/Avx512BW (net471 has no intrinsics -> false), so
it auto-engages on the VNNI serving hardware and stays on fp32 otherwise. weights are
still read + kept as Q8_0 regardless (RAM benefit); only the forward path is gated.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…arness

decode profiling (new DEVHOST_DECODE mode: real incremental paged generation,
reports tok/s) showed ComputeTiledPagedAttention read the KV history one locked
ReadKey/ReadValue per position -> O(seqLen) contended Monitor.Enter per layer per
token. add PagedKVCache.ReadKeyValueRange (one lock, whole range) and read the
layer's KV once into a pooled buffer; the attention loop then indexes it lock-free.
removes the per-position locking (helps concurrent multi-sequence load where the
cache lock is genuinely contended). 14/14 paged-attention + incremental-generation
tests green (identical output). single-stream decode ~130ms/token is serial
per-token compute-bound (POOL_THREADS 1/2/4 all equal), not lock-bound — the
forward-thread lock samples were blocked-wait, not CPU.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
MatVecMul (the q/k/v/o + attention decode projections, run 7x per layer per token
on the incremental path) was a pure scalar double loop -- ~150M scalar FMAs/token
for a 135M model, the dominant fixed per-token decode cost. vectorize the inner dot
with System.Numerics.Vector<float> (8-wide FMA + horizontal sum + scalar tail),
net471 keeps the scalar path. decode 6.9 -> 11.3 tok/s (145 -> 88 ms/token, 1.6x)
on smollm2-135m. 14/14 paged + incremental tests green (identical output).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@vercel

vercel Bot commented Jul 21, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
aidotnet_website Ready Ready Preview, Comment Jul 21, 2026 1:04am
aidotnet-playground-api Ready Ready Preview, Comment Jul 21, 2026 1:04am

@coderabbitai

coderabbitai Bot commented Jul 21, 2026 •

Copy link
Copy Markdown
Contributor

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: c0d90642-4208-40cb-83c3-2afc59d8edd9

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/device-agnostic-model

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown
Contributor

Commit messages auto-fixed

One or more commit messages did not follow Conventional Commits, so they were rewritten to comply (subject case, header length ≤ 100, valid type). Each commit and its diff were preserved — no squashing.

The branch was force-pushed with the corrected messages. If you have local work on this branch, run git pull --rebase (or reset to the remote) before pushing again.

@ooples
ooples force-pushed the feat/device-agnostic-model branch from b9e87df to 284d448 Compare July 21, 2026 01:03
@ooples
ooples merged commit 5b89135 into feat/serving-integration Jul 21, 2026
11 of 16 checks passed
@ooples
ooples deleted the feat/device-agnostic-model branch July 21, 2026 03:17
ooples added a commit that referenced this pull request Jul 21, 2026
…ecode/prefill perf (#1917)

* feat(device): model.To(DeviceInfo) / layer.To(DeviceInfo) placement API

PyTorch-style device placement, enum-only (no magic-string overloads). LayerBase.To
moves registered parameters + buffers and recurses sublayers; NeuralNetworkBase.To
moves every layer. Skips zero-length (deferred) params. First user-facing piece of
the device-agnostic execution model; builds net10.0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(device): facade ConfigureModel(source, DeviceInfo) overload (Level 1)

Load a pretrained checkpoint and place the whole model on a device in one call —
the common "load onto my GPU" case — via the type-safe enum DeviceInfo (no device
strings). Additive overload: existing ConfigureModel(source) callers are unchanged.
Delegates to model.To(device). Builds net10.0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(device): regression test for model.To(OpenCL()) placement -> 7042

Proves the PyTorch-style placement API (model.To(DeviceInfo.OpenCL())) drives
correct GPU execution to llama.cpp's greedy token. Uses AutoDetectAndConfigureGpu
so placement (Tensor.To -> global backend) and execution (Current dispatcher)
share one backend; skips when no GPU is wired in-process (validates on DevHost/CUDA).
Surfaced that To() and Current use different backend acquisition — unify in Phase 1.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(device): phase-2 fusion acceptance test (skipped) + document the gap

Runs the whole decoder forward inside a DeferredScope (BeginDeferredScope ->
Predict -> Execute) as the acceptance check for transparent fusion. Today it
returns token 0 (all-zero logits) instead of 7042: the graph capture/replay does
not materialize the decoder's output, so transparent fusion needs the graph path
(output binding + RoPE/GQA-SDPA/RMSNorm recording) debugged before it can wrap
Predict. Skipped with that reason; un-skip when the graph path is fixed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(serving): inferenceOptimizer recurses into PreLNTransformerBlock for GQA KV-cache

Adds PreLNTransformerBlock.ReplaceAttention (mirrors TransformerEncoderBlock) and a
PreLNTransformerBlock case in ApplyAttentionOptimizations that swaps the block's nested
GroupedQueryAttentionLayer for a KV-cached CachedGroupedQueryAttention (shared helper
BuildCachedGqaReplacement, reused by the top-level GQA case). This is the first piece of
wiring the GGUF/LLaMA decoder (GQA nested in PreLNTransformerBlock) to incremental KV-cache
decode. Still inert end-to-end: InitializeGQAKVCache's collection scan is also top-level
(won't yet find the nested cached GQA), and ServableModelWrapper requires a paged cache
that GQA lacks — so the incremental clone is still discarded and serving falls back to the
eager model (no regression). Remaining: recurse InitializeGQAKVCache; accept the non-paged
GQA cache as the incremental cache. Builds net10.0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(serving): enumerateAttentionHosts reaches PreLNTransformerBlock attention

The shared attention-host enumerator (used by the GQA/paged KV-cache init,
attention-optimizability detection, and quantization scans) recursed into
TransformerEncoderBlock/TransformerDecoderBlock but not PreLNTransformerBlock,
where LLaMA/GGUF decoders host their grouped-query attention. Add that case so
InitializeGQAKVCache (and the other scans) find the nested CachedGroupedQueryAttention.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(serving): make PagedAttentionKernel grouped-query-attention aware

Add PagedAttentionConfig.NumQueryHeads (0 => same as NumHeads = MHA). NumHeads is
now the KV/cache head count; query heads may outnumber it under GQA, with each KV
head shared by NumQueryHeads/NumHeads query heads (kvHead = qHead / group). Threaded
through every compute path — ComputeAttention, ComputeTiledPagedAttention (decode),
ComputeContiguousCausalPrefill (prefill), ComputeBatchedAttention, and the fused
Forward/ForwardQuantized (asymmetric q_proj vs k_proj/v_proj, matching HF). K/V
buffers size to the KV heads the cache stores; per-query-head online-softmax scratch.
MHA stays the group==1 special case (all 43 PagedAttention tests unchanged).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(serving): make PagedCachedMultiHeadAttention grouped-query-attention aware

Add kvHeadCount ctor param (0 => headCount = MHA). K/V weights are now
[embDim, kvHeadCount*headDim] (narrower under GQA); Q/O stay [embDim, embDim].
Threaded kvProjDim through every projection + cache write in the per-token,
batched-GEMM, and contiguous-prefill paths; RoPE applies to Q over query heads
and K over KV heads; param serialization and int8 quant use the real per-weight
widths. Expose KVHeadCount for the optimizer to build the paged config/kernel.

Head split/merge and repeat_kv now use vectorized library ops (Reshape/Transpose
and Engine.TensorGather on the head axis) instead of scalar nested loops. MHA is
the kvHeadCount==headCount special case; all 43 PagedAttention tests unchanged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(serving): wire grouped-query decoders onto the paged KV-cache path

InitializePagedKVCache now sizes the PagedKVCache by the KV-head count and passes
NumQueryHeads to the kernel so it repeats each KV head across its query-head group.
BuildPagedGqaReplacement converts a GroupedQueryAttentionLayer (top-level or nested
in PreLNTransformerBlock) to a PagedCachedMultiHeadAttention with the source KV-head
count, copying its [Q][K][V][O][outBias] parameters and RoPE/ALiBi config; the GQA
branches prefer it when paged KV is enabled. Guarded by GroupedQueryAttentionLayer.
UsesProjectionBias: models with Q/K/V projection bias (Qwen2-style) fall back to the
contiguous CachedGroupedQueryAttention, which the paged layer cannot represent.

200 paged/optimizer/KVCache/GQA tests green net10.0; both TFMs build.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(serving): paged grouped-query attention matches the original decoder

Numerical-parity test: a tiny GQA decoder (numHeads=4, numKVHeads=2, RoPE) nested
in a PreLNTransformerBlock is optimized with the paged KV cache enabled, and the
optimized forward matches the original within 1e-3 — proving the paged-GQA kernel
repeat-KV, the layer's narrow K/V projections, the [Q][K][V][O][outBias] weight
copy, and the interleaved-RoPE convention are all faithful. Also asserts the GQA is
rewritten to PagedCachedMultiHeadAttention and the paged KV cache is live.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(serialize): groupedQueryAttentionLayer round-trips RoPE, mask, bias, softcap

A cloned GQA layer (serialize -> deserialize) silently lost its RoPE positional
encoding, causal mask, custom head dimension, Q/K/V projection bias, and attention
logit soft-cap: GetMetadata never persisted them and the deserializer defaulted the
bools to false with no ConfigurePositionalEncoding call. So a cloned decoder computed
bidirectional, RoPE-less attention and diverged — which is exactly what breaks the
incremental-serving clone of a GGUF/LLaMA decoder. Persist all of them in GetMetadata
and restore them (incl. RoPE via the layer's ConfigurePositionalEncoding) on deserialize.

New test clones a tiny RoPE+causal GQA model and asserts forward parity within 1e-4.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(serialize): make PreLNTransformerBlock cloneable so paged serving engages

The block hosts a polymorphic (T5/MHA/GQA) attention sublayer, so it had no
deserialization constructor and cloning threw — which silently disabled the paged
incremental-generation clone (ServableModelWrapper.BuildIncrementalModel) for every
GGUF/LLaMA decoder. GetMetadata now persists the block dims + FFN activation + the
nested attention as an 'Attn.'-prefixed self-contained sub-blob (type + its metadata
+ shapes); a new DeserializationHelper branch rebuilds the attention recursively via
CreateLayerFromType (E1 keeps its RoPE/mask) and constructs the block. OnFirstForward
resolves the lazy norms + FFN in forward order (deser/resolve path only, so no RNG
perturbation on a normal forward) so ParameterCount is correct before SetParameters.

The paged-GQA parity test now runs with cloneModel: true (the real serving path) and
still matches the original within 1e-3.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(serving): grouped-query decoder builds the paged incremental path

End-to-end proof that paged-GQA engages in serving: a numHeads>numKVHeads PreLN
decoder (RoPE) wrapped in ServableModelWrapper now reports SupportsIncrementalGeneration
true and generates in-range tokens. Exercises the full chain — GQA-aware paged kernel,
narrow-K/V paged attention layer, optimizer GQA->paged rewrite, and PreLNTransformerBlock
+ GQA serialization (the clone BuildIncrementalModel performs). Before this the clone
threw and every LLaMA/GGUF-style decoder silently fell back to the stateless path.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(inference): cachedGroupedQueryAttention forward onto recordable engine ops [prototype]

Rewrites the inference-only cached GQA layer to apply RoPE via Engine.ApplyRoPEInterleaved
and grouped-query attention via Engine.ScaledDotProductAttentionGqa (unexpanded K/V) — a
device-agnostic, GPU-graph-recordable forward that drops the two managed unrecordable
hotspots (managed RoPE + ExpandKVHeads). RotaryPositionalEncodingLayer exposes its cos/sin
caches (GetInterleavedCaches) for the fused op. ALiBi keeps the expand + FlashAttention path.

Verified: the optimized (non-paged) model's forward matches the original managed GQA decoder
within 1e-3 (InferenceOptimizer_CachedGroupedQueryAttention_RewrittenForward_MatchesOriginal),
run against a local AiDotNet.Tensors probe.

PROTOTYPE / RELEASE-GATED: the pin is a LOCAL probe (0.117.200-devexec) of the unreleased
AiDotNet.Tensors PR #828 (device-agnostic execution). It MUST be bumped to the released
AiDotNet.Tensors version before this lands / any AiDotNet PR — CI cannot restore the probe.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* chore: bump probe pin to 0.117.201-devexec (GQA-SDPA causal KV-cache offset fix)

Local probe of the Tensors branch after the OpenCL GQA-SDPA causal offset fix. Still
release-gated — bumps to the released AiDotNet.Tensors before any AiDotNet PR.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(gguf): serve Q8_0 weights quantized via block-Q8_0 GEMM (decode) + profiling harness

keeps GGUF Q8_0 linear weights in their native int8 layout instead of expanding to
fp32 at load, and runs the block-Q8_0 GEMM (Tensors Q8BlockGemm, ggml
q8_0 parity) directly on them for float inference. GgufFile.TryReadQ8_0Raw reads
the native blocks (int8 + per-32 fp16 scale); GgufModelSource.TryReadQ8_0 maps the
HF name; LlamaModelBuilder.LoadDense installs them on the FFN + lm_head DenseLayers;
DenseLayer runs the quantized GEMM on the float inference path. gated to small M
(decode/small-batch, where the naive int8 kernel is bandwidth-bound and beats fp32
BLAS); large-M prefill stays on fp32 until the int8 GEMM is register-tiled, so
prefill is unchanged (451 tok/s) while decode runs on int8. gguf 7042 llama.cpp
parity preserved. also adds the DEVHOST_PROFILE steady-state loop used to profile
the forward. NOTE: release-gated on Tensors Q8BlockGemm (probe pin 0.118.10-q8).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* perf(gguf): gate Q8_0 quantized forward on VNNI hardware

the block-Q8_0 GEMM only beats the tuned fp32 BLAS when the CPU has a 1-instruction
VNNI int8 multiply-accumulate (AvxVnni / Avx512 vpdpbusd). on AVX2-only CPUs the
3-instruction maddubs dot is no faster, so engaging it there would regress. gate the
DenseLayer quantized path on AvxVnni/Avx512BW (net471 has no intrinsics -> false), so
it auto-engages on the VNNI serving hardware and stays on fp32 otherwise. weights are
still read + kept as Q8_0 regardless (RAM benefit); only the forward path is gated.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* perf(serving): bulk-read paged KV under one lock + decode benchmark harness

decode profiling (new DEVHOST_DECODE mode: real incremental paged generation,
reports tok/s) showed ComputeTiledPagedAttention read the KV history one locked
ReadKey/ReadValue per position -> O(seqLen) contended Monitor.Enter per layer per
token. add PagedKVCache.ReadKeyValueRange (one lock, whole range) and read the
layer's KV once into a pooled buffer; the attention loop then indexes it lock-free.
removes the per-position locking (helps concurrent multi-sequence load where the
cache lock is genuinely contended). 14/14 paged-attention + incremental-generation
tests green (identical output). single-stream decode ~130ms/token is serial
per-token compute-bound (POOL_THREADS 1/2/4 all equal), not lock-bound — the
forward-thread lock samples were blocked-wait, not CPU.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* perf(decode): vectorize the per-token projection matvec (scalar -> AVX2)

MatVecMul (the q/k/v/o + attention decode projections, run 7x per layer per token
on the incremental path) was a pure scalar double loop -- ~150M scalar FMAs/token
for a 135M model, the dominant fixed per-token decode cost. vectorize the inner dot
with System.Numerics.Vector<float> (8-wide FMA + horizontal sum + scalar tail),
net471 keeps the scalar path. decode 6.9 -> 11.3 tok/s (145 -> 88 ms/token, 1.6x)
on smollm2-135m. 14/14 paged + incremental tests green (identical output).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

This branch was successfully deployed

2 active deployments
Preview – aidotnet-playground-api — 284d448c Deployed Jul 21, 2026 by vercel[bot]
Preview – aidotnet_website — 284d448c Deployed Jul 21, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant