feat: device-agnostic execution + GGUF Q8_0 quantized serving + CPU decode/prefill perf - #1917
Conversation
PyTorch-style device placement, enum-only (no magic-string overloads). LayerBase.To moves registered parameters + buffers and recurses sublayers; NeuralNetworkBase.To moves every layer. Skips zero-length (deferred) params. First user-facing piece of the device-agnostic execution model; builds net10.0. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…el 1) Load a pretrained checkpoint and place the whole model on a device in one call — the common "load onto my GPU" case — via the type-safe enum DeviceInfo (no device strings). Additive overload: existing ConfigureModel(source) callers are unchanged. Delegates to model.To(device). Builds net10.0. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Proves the PyTorch-style placement API (model.To(DeviceInfo.OpenCL())) drives correct GPU execution to llama.cpp's greedy token. Uses AutoDetectAndConfigureGpu so placement (Tensor.To -> global backend) and execution (Current dispatcher) share one backend; skips when no GPU is wired in-process (validates on DevHost/CUDA). Surfaced that To() and Current use different backend acquisition — unify in Phase 1. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
… gap Runs the whole decoder forward inside a DeferredScope (BeginDeferredScope -> Predict -> Execute) as the acceptance check for transparent fusion. Today it returns token 0 (all-zero logits) instead of 7042: the graph capture/replay does not materialize the decoder's output, so transparent fusion needs the graph path (output binding + RoPE/GQA-SDPA/RMSNorm recording) debugged before it can wrap Predict. Skipped with that reason; un-skip when the graph path is fixed. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
… for GQA KV-cache Adds PreLNTransformerBlock.ReplaceAttention (mirrors TransformerEncoderBlock) and a PreLNTransformerBlock case in ApplyAttentionOptimizations that swaps the block's nested GroupedQueryAttentionLayer for a KV-cached CachedGroupedQueryAttention (shared helper BuildCachedGqaReplacement, reused by the top-level GQA case). This is the first piece of wiring the GGUF/LLaMA decoder (GQA nested in PreLNTransformerBlock) to incremental KV-cache decode. Still inert end-to-end: InitializeGQAKVCache's collection scan is also top-level (won't yet find the nested cached GQA), and ServableModelWrapper requires a paged cache that GQA lacks — so the incremental clone is still discarded and serving falls back to the eager model (no regression). Remaining: recurse InitializeGQAKVCache; accept the non-paged GQA cache as the incremental cache. Builds net10.0. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…attention The shared attention-host enumerator (used by the GQA/paged KV-cache init, attention-optimizability detection, and quantization scans) recursed into TransformerEncoderBlock/TransformerDecoderBlock but not PreLNTransformerBlock, where LLaMA/GGUF decoders host their grouped-query attention. Add that case so InitializeGQAKVCache (and the other scans) find the nested CachedGroupedQueryAttention. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add PagedAttentionConfig.NumQueryHeads (0 => same as NumHeads = MHA). NumHeads is now the KV/cache head count; query heads may outnumber it under GQA, with each KV head shared by NumQueryHeads/NumHeads query heads (kvHead = qHead / group). Threaded through every compute path — ComputeAttention, ComputeTiledPagedAttention (decode), ComputeContiguousCausalPrefill (prefill), ComputeBatchedAttention, and the fused Forward/ForwardQuantized (asymmetric q_proj vs k_proj/v_proj, matching HF). K/V buffers size to the KV heads the cache stores; per-query-head online-softmax scratch. MHA stays the group==1 special case (all 43 PagedAttention tests unchanged). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…tion aware Add kvHeadCount ctor param (0 => headCount = MHA). K/V weights are now [embDim, kvHeadCount*headDim] (narrower under GQA); Q/O stay [embDim, embDim]. Threaded kvProjDim through every projection + cache write in the per-token, batched-GEMM, and contiguous-prefill paths; RoPE applies to Q over query heads and K over KV heads; param serialization and int8 quant use the real per-weight widths. Expose KVHeadCount for the optimizer to build the paged config/kernel. Head split/merge and repeat_kv now use vectorized library ops (Reshape/Transpose and Engine.TensorGather on the head axis) instead of scalar nested loops. MHA is the kvHeadCount==headCount special case; all 43 PagedAttention tests unchanged. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
InitializePagedKVCache now sizes the PagedKVCache by the KV-head count and passes NumQueryHeads to the kernel so it repeats each KV head across its query-head group. BuildPagedGqaReplacement converts a GroupedQueryAttentionLayer (top-level or nested in PreLNTransformerBlock) to a PagedCachedMultiHeadAttention with the source KV-head count, copying its [Q][K][V][O][outBias] parameters and RoPE/ALiBi config; the GQA branches prefer it when paged KV is enabled. Guarded by GroupedQueryAttentionLayer. UsesProjectionBias: models with Q/K/V projection bias (Qwen2-style) fall back to the contiguous CachedGroupedQueryAttention, which the paged layer cannot represent. 200 paged/optimizer/KVCache/GQA tests green net10.0; both TFMs build. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…oder Numerical-parity test: a tiny GQA decoder (numHeads=4, numKVHeads=2, RoPE) nested in a PreLNTransformerBlock is optimized with the paged KV cache enabled, and the optimized forward matches the original within 1e-3 — proving the paged-GQA kernel repeat-KV, the layer's narrow K/V projections, the [Q][K][V][O][outBias] weight copy, and the interleaved-RoPE convention are all faithful. Also asserts the GQA is rewritten to PagedCachedMultiHeadAttention and the paged KV cache is live. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…as, softcap A cloned GQA layer (serialize -> deserialize) silently lost its RoPE positional encoding, causal mask, custom head dimension, Q/K/V projection bias, and attention logit soft-cap: GetMetadata never persisted them and the deserializer defaulted the bools to false with no ConfigurePositionalEncoding call. So a cloned decoder computed bidirectional, RoPE-less attention and diverged — which is exactly what breaks the incremental-serving clone of a GGUF/LLaMA decoder. Persist all of them in GetMetadata and restore them (incl. RoPE via the layer's ConfigurePositionalEncoding) on deserialize. New test clones a tiny RoPE+causal GQA model and asserts forward parity within 1e-4. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…g engages The block hosts a polymorphic (T5/MHA/GQA) attention sublayer, so it had no deserialization constructor and cloning threw — which silently disabled the paged incremental-generation clone (ServableModelWrapper.BuildIncrementalModel) for every GGUF/LLaMA decoder. GetMetadata now persists the block dims + FFN activation + the nested attention as an 'Attn.'-prefixed self-contained sub-blob (type + its metadata + shapes); a new DeserializationHelper branch rebuilds the attention recursively via CreateLayerFromType (E1 keeps its RoPE/mask) and constructs the block. OnFirstForward resolves the lazy norms + FFN in forward order (deser/resolve path only, so no RNG perturbation on a normal forward) so ParameterCount is correct before SetParameters. The paged-GQA parity test now runs with cloneModel: true (the real serving path) and still matches the original within 1e-3. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
End-to-end proof that paged-GQA engages in serving: a numHeads>numKVHeads PreLN decoder (RoPE) wrapped in ServableModelWrapper now reports SupportsIncrementalGeneration true and generates in-range tokens. Exercises the full chain — GQA-aware paged kernel, narrow-K/V paged attention layer, optimizer GQA->paged rewrite, and PreLNTransformerBlock + GQA serialization (the clone BuildIncrementalModel performs). Before this the clone threw and every LLaMA/GGUF-style decoder silently fell back to the stateless path. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…engine ops [prototype] Rewrites the inference-only cached GQA layer to apply RoPE via Engine.ApplyRoPEInterleaved and grouped-query attention via Engine.ScaledDotProductAttentionGqa (unexpanded K/V) — a device-agnostic, GPU-graph-recordable forward that drops the two managed unrecordable hotspots (managed RoPE + ExpandKVHeads). RotaryPositionalEncodingLayer exposes its cos/sin caches (GetInterleavedCaches) for the fused op. ALiBi keeps the expand + FlashAttention path. Verified: the optimized (non-paged) model's forward matches the original managed GQA decoder within 1e-3 (InferenceOptimizer_CachedGroupedQueryAttention_RewrittenForward_MatchesOriginal), run against a local AiDotNet.Tensors probe. PROTOTYPE / RELEASE-GATED: the pin is a LOCAL probe (0.117.200-devexec) of the unreleased AiDotNet.Tensors PR #828 (device-agnostic execution). It MUST be bumped to the released AiDotNet.Tensors version before this lands / any AiDotNet PR — CI cannot restore the probe. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…offset fix) Local probe of the Tensors branch after the OpenCL GQA-SDPA causal offset fix. Still release-gated — bumps to the released AiDotNet.Tensors before any AiDotNet PR. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
… + profiling harness keeps GGUF Q8_0 linear weights in their native int8 layout instead of expanding to fp32 at load, and runs the block-Q8_0 GEMM (Tensors Q8BlockGemm, ggml q8_0 parity) directly on them for float inference. GgufFile.TryReadQ8_0Raw reads the native blocks (int8 + per-32 fp16 scale); GgufModelSource.TryReadQ8_0 maps the HF name; LlamaModelBuilder.LoadDense installs them on the FFN + lm_head DenseLayers; DenseLayer runs the quantized GEMM on the float inference path. gated to small M (decode/small-batch, where the naive int8 kernel is bandwidth-bound and beats fp32 BLAS); large-M prefill stays on fp32 until the int8 GEMM is register-tiled, so prefill is unchanged (451 tok/s) while decode runs on int8. gguf 7042 llama.cpp parity preserved. also adds the DEVHOST_PROFILE steady-state loop used to profile the forward. NOTE: release-gated on Tensors Q8BlockGemm (probe pin 0.118.10-q8). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
the block-Q8_0 GEMM only beats the tuned fp32 BLAS when the CPU has a 1-instruction VNNI int8 multiply-accumulate (AvxVnni / Avx512 vpdpbusd). on AVX2-only CPUs the 3-instruction maddubs dot is no faster, so engaging it there would regress. gate the DenseLayer quantized path on AvxVnni/Avx512BW (net471 has no intrinsics -> false), so it auto-engages on the VNNI serving hardware and stays on fp32 otherwise. weights are still read + kept as Q8_0 regardless (RAM benefit); only the forward path is gated. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…arness decode profiling (new DEVHOST_DECODE mode: real incremental paged generation, reports tok/s) showed ComputeTiledPagedAttention read the KV history one locked ReadKey/ReadValue per position -> O(seqLen) contended Monitor.Enter per layer per token. add PagedKVCache.ReadKeyValueRange (one lock, whole range) and read the layer's KV once into a pooled buffer; the attention loop then indexes it lock-free. removes the per-position locking (helps concurrent multi-sequence load where the cache lock is genuinely contended). 14/14 paged-attention + incremental-generation tests green (identical output). single-stream decode ~130ms/token is serial per-token compute-bound (POOL_THREADS 1/2/4 all equal), not lock-bound — the forward-thread lock samples were blocked-wait, not CPU. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
MatVecMul (the q/k/v/o + attention decode projections, run 7x per layer per token on the incremental path) was a pure scalar double loop -- ~150M scalar FMAs/token for a 135M model, the dominant fixed per-token decode cost. vectorize the inner dot with System.Numerics.Vector<float> (8-wide FMA + horizontal sum + scalar tail), net471 keeps the scalar path. decode 6.9 -> 11.3 tok/s (145 -> 88 ms/token, 1.6x) on smollm2-135m. 14/14 paged + incremental tests green (identical output). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
|
Important Review skippedAuto reviews are disabled on base/target branches other than the default branch. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Pro Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
Commit messages auto-fixedOne or more commit messages did not follow Conventional Commits, so they were rewritten to comply (subject case, header length ≤ 100, valid type). Each commit and its diff were preserved — no squashing. The branch was force-pushed with the corrected messages. If you have local work on this branch, run |
b9e87df to
284d448
Compare
…ecode/prefill perf (#1917) * feat(device): model.To(DeviceInfo) / layer.To(DeviceInfo) placement API PyTorch-style device placement, enum-only (no magic-string overloads). LayerBase.To moves registered parameters + buffers and recurses sublayers; NeuralNetworkBase.To moves every layer. Skips zero-length (deferred) params. First user-facing piece of the device-agnostic execution model; builds net10.0. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(device): facade ConfigureModel(source, DeviceInfo) overload (Level 1) Load a pretrained checkpoint and place the whole model on a device in one call — the common "load onto my GPU" case — via the type-safe enum DeviceInfo (no device strings). Additive overload: existing ConfigureModel(source) callers are unchanged. Delegates to model.To(device). Builds net10.0. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(device): regression test for model.To(OpenCL()) placement -> 7042 Proves the PyTorch-style placement API (model.To(DeviceInfo.OpenCL())) drives correct GPU execution to llama.cpp's greedy token. Uses AutoDetectAndConfigureGpu so placement (Tensor.To -> global backend) and execution (Current dispatcher) share one backend; skips when no GPU is wired in-process (validates on DevHost/CUDA). Surfaced that To() and Current use different backend acquisition — unify in Phase 1. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(device): phase-2 fusion acceptance test (skipped) + document the gap Runs the whole decoder forward inside a DeferredScope (BeginDeferredScope -> Predict -> Execute) as the acceptance check for transparent fusion. Today it returns token 0 (all-zero logits) instead of 7042: the graph capture/replay does not materialize the decoder's output, so transparent fusion needs the graph path (output binding + RoPE/GQA-SDPA/RMSNorm recording) debugged before it can wrap Predict. Skipped with that reason; un-skip when the graph path is fixed. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(serving): inferenceOptimizer recurses into PreLNTransformerBlock for GQA KV-cache Adds PreLNTransformerBlock.ReplaceAttention (mirrors TransformerEncoderBlock) and a PreLNTransformerBlock case in ApplyAttentionOptimizations that swaps the block's nested GroupedQueryAttentionLayer for a KV-cached CachedGroupedQueryAttention (shared helper BuildCachedGqaReplacement, reused by the top-level GQA case). This is the first piece of wiring the GGUF/LLaMA decoder (GQA nested in PreLNTransformerBlock) to incremental KV-cache decode. Still inert end-to-end: InitializeGQAKVCache's collection scan is also top-level (won't yet find the nested cached GQA), and ServableModelWrapper requires a paged cache that GQA lacks — so the incremental clone is still discarded and serving falls back to the eager model (no regression). Remaining: recurse InitializeGQAKVCache; accept the non-paged GQA cache as the incremental cache. Builds net10.0. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(serving): enumerateAttentionHosts reaches PreLNTransformerBlock attention The shared attention-host enumerator (used by the GQA/paged KV-cache init, attention-optimizability detection, and quantization scans) recursed into TransformerEncoderBlock/TransformerDecoderBlock but not PreLNTransformerBlock, where LLaMA/GGUF decoders host their grouped-query attention. Add that case so InitializeGQAKVCache (and the other scans) find the nested CachedGroupedQueryAttention. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(serving): make PagedAttentionKernel grouped-query-attention aware Add PagedAttentionConfig.NumQueryHeads (0 => same as NumHeads = MHA). NumHeads is now the KV/cache head count; query heads may outnumber it under GQA, with each KV head shared by NumQueryHeads/NumHeads query heads (kvHead = qHead / group). Threaded through every compute path — ComputeAttention, ComputeTiledPagedAttention (decode), ComputeContiguousCausalPrefill (prefill), ComputeBatchedAttention, and the fused Forward/ForwardQuantized (asymmetric q_proj vs k_proj/v_proj, matching HF). K/V buffers size to the KV heads the cache stores; per-query-head online-softmax scratch. MHA stays the group==1 special case (all 43 PagedAttention tests unchanged). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(serving): make PagedCachedMultiHeadAttention grouped-query-attention aware Add kvHeadCount ctor param (0 => headCount = MHA). K/V weights are now [embDim, kvHeadCount*headDim] (narrower under GQA); Q/O stay [embDim, embDim]. Threaded kvProjDim through every projection + cache write in the per-token, batched-GEMM, and contiguous-prefill paths; RoPE applies to Q over query heads and K over KV heads; param serialization and int8 quant use the real per-weight widths. Expose KVHeadCount for the optimizer to build the paged config/kernel. Head split/merge and repeat_kv now use vectorized library ops (Reshape/Transpose and Engine.TensorGather on the head axis) instead of scalar nested loops. MHA is the kvHeadCount==headCount special case; all 43 PagedAttention tests unchanged. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(serving): wire grouped-query decoders onto the paged KV-cache path InitializePagedKVCache now sizes the PagedKVCache by the KV-head count and passes NumQueryHeads to the kernel so it repeats each KV head across its query-head group. BuildPagedGqaReplacement converts a GroupedQueryAttentionLayer (top-level or nested in PreLNTransformerBlock) to a PagedCachedMultiHeadAttention with the source KV-head count, copying its [Q][K][V][O][outBias] parameters and RoPE/ALiBi config; the GQA branches prefer it when paged KV is enabled. Guarded by GroupedQueryAttentionLayer. UsesProjectionBias: models with Q/K/V projection bias (Qwen2-style) fall back to the contiguous CachedGroupedQueryAttention, which the paged layer cannot represent. 200 paged/optimizer/KVCache/GQA tests green net10.0; both TFMs build. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(serving): paged grouped-query attention matches the original decoder Numerical-parity test: a tiny GQA decoder (numHeads=4, numKVHeads=2, RoPE) nested in a PreLNTransformerBlock is optimized with the paged KV cache enabled, and the optimized forward matches the original within 1e-3 — proving the paged-GQA kernel repeat-KV, the layer's narrow K/V projections, the [Q][K][V][O][outBias] weight copy, and the interleaved-RoPE convention are all faithful. Also asserts the GQA is rewritten to PagedCachedMultiHeadAttention and the paged KV cache is live. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(serialize): groupedQueryAttentionLayer round-trips RoPE, mask, bias, softcap A cloned GQA layer (serialize -> deserialize) silently lost its RoPE positional encoding, causal mask, custom head dimension, Q/K/V projection bias, and attention logit soft-cap: GetMetadata never persisted them and the deserializer defaulted the bools to false with no ConfigurePositionalEncoding call. So a cloned decoder computed bidirectional, RoPE-less attention and diverged — which is exactly what breaks the incremental-serving clone of a GGUF/LLaMA decoder. Persist all of them in GetMetadata and restore them (incl. RoPE via the layer's ConfigurePositionalEncoding) on deserialize. New test clones a tiny RoPE+causal GQA model and asserts forward parity within 1e-4. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(serialize): make PreLNTransformerBlock cloneable so paged serving engages The block hosts a polymorphic (T5/MHA/GQA) attention sublayer, so it had no deserialization constructor and cloning threw — which silently disabled the paged incremental-generation clone (ServableModelWrapper.BuildIncrementalModel) for every GGUF/LLaMA decoder. GetMetadata now persists the block dims + FFN activation + the nested attention as an 'Attn.'-prefixed self-contained sub-blob (type + its metadata + shapes); a new DeserializationHelper branch rebuilds the attention recursively via CreateLayerFromType (E1 keeps its RoPE/mask) and constructs the block. OnFirstForward resolves the lazy norms + FFN in forward order (deser/resolve path only, so no RNG perturbation on a normal forward) so ParameterCount is correct before SetParameters. The paged-GQA parity test now runs with cloneModel: true (the real serving path) and still matches the original within 1e-3. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(serving): grouped-query decoder builds the paged incremental path End-to-end proof that paged-GQA engages in serving: a numHeads>numKVHeads PreLN decoder (RoPE) wrapped in ServableModelWrapper now reports SupportsIncrementalGeneration true and generates in-range tokens. Exercises the full chain — GQA-aware paged kernel, narrow-K/V paged attention layer, optimizer GQA->paged rewrite, and PreLNTransformerBlock + GQA serialization (the clone BuildIncrementalModel performs). Before this the clone threw and every LLaMA/GGUF-style decoder silently fell back to the stateless path. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(inference): cachedGroupedQueryAttention forward onto recordable engine ops [prototype] Rewrites the inference-only cached GQA layer to apply RoPE via Engine.ApplyRoPEInterleaved and grouped-query attention via Engine.ScaledDotProductAttentionGqa (unexpanded K/V) — a device-agnostic, GPU-graph-recordable forward that drops the two managed unrecordable hotspots (managed RoPE + ExpandKVHeads). RotaryPositionalEncodingLayer exposes its cos/sin caches (GetInterleavedCaches) for the fused op. ALiBi keeps the expand + FlashAttention path. Verified: the optimized (non-paged) model's forward matches the original managed GQA decoder within 1e-3 (InferenceOptimizer_CachedGroupedQueryAttention_RewrittenForward_MatchesOriginal), run against a local AiDotNet.Tensors probe. PROTOTYPE / RELEASE-GATED: the pin is a LOCAL probe (0.117.200-devexec) of the unreleased AiDotNet.Tensors PR #828 (device-agnostic execution). It MUST be bumped to the released AiDotNet.Tensors version before this lands / any AiDotNet PR — CI cannot restore the probe. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore: bump probe pin to 0.117.201-devexec (GQA-SDPA causal KV-cache offset fix) Local probe of the Tensors branch after the OpenCL GQA-SDPA causal offset fix. Still release-gated — bumps to the released AiDotNet.Tensors before any AiDotNet PR. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(gguf): serve Q8_0 weights quantized via block-Q8_0 GEMM (decode) + profiling harness keeps GGUF Q8_0 linear weights in their native int8 layout instead of expanding to fp32 at load, and runs the block-Q8_0 GEMM (Tensors Q8BlockGemm, ggml q8_0 parity) directly on them for float inference. GgufFile.TryReadQ8_0Raw reads the native blocks (int8 + per-32 fp16 scale); GgufModelSource.TryReadQ8_0 maps the HF name; LlamaModelBuilder.LoadDense installs them on the FFN + lm_head DenseLayers; DenseLayer runs the quantized GEMM on the float inference path. gated to small M (decode/small-batch, where the naive int8 kernel is bandwidth-bound and beats fp32 BLAS); large-M prefill stays on fp32 until the int8 GEMM is register-tiled, so prefill is unchanged (451 tok/s) while decode runs on int8. gguf 7042 llama.cpp parity preserved. also adds the DEVHOST_PROFILE steady-state loop used to profile the forward. NOTE: release-gated on Tensors Q8BlockGemm (probe pin 0.118.10-q8). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * perf(gguf): gate Q8_0 quantized forward on VNNI hardware the block-Q8_0 GEMM only beats the tuned fp32 BLAS when the CPU has a 1-instruction VNNI int8 multiply-accumulate (AvxVnni / Avx512 vpdpbusd). on AVX2-only CPUs the 3-instruction maddubs dot is no faster, so engaging it there would regress. gate the DenseLayer quantized path on AvxVnni/Avx512BW (net471 has no intrinsics -> false), so it auto-engages on the VNNI serving hardware and stays on fp32 otherwise. weights are still read + kept as Q8_0 regardless (RAM benefit); only the forward path is gated. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * perf(serving): bulk-read paged KV under one lock + decode benchmark harness decode profiling (new DEVHOST_DECODE mode: real incremental paged generation, reports tok/s) showed ComputeTiledPagedAttention read the KV history one locked ReadKey/ReadValue per position -> O(seqLen) contended Monitor.Enter per layer per token. add PagedKVCache.ReadKeyValueRange (one lock, whole range) and read the layer's KV once into a pooled buffer; the attention loop then indexes it lock-free. removes the per-position locking (helps concurrent multi-sequence load where the cache lock is genuinely contended). 14/14 paged-attention + incremental-generation tests green (identical output). single-stream decode ~130ms/token is serial per-token compute-bound (POOL_THREADS 1/2/4 all equal), not lock-bound — the forward-thread lock samples were blocked-wait, not CPU. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * perf(decode): vectorize the per-token projection matvec (scalar -> AVX2) MatVecMul (the q/k/v/o + attention decode projections, run 7x per layer per token on the incremental path) was a pure scalar double loop -- ~150M scalar FMAs/token for a 135M model, the dominant fixed per-token decode cost. vectorize the inner dot with System.Numerics.Vector<float> (8-wide FMA + horizontal sum + scalar tail), net471 keeps the scalar path. decode 6.9 -> 11.3 tok/s (145 -> 88 ms/token, 1.6x) on smollm2-135m. 14/14 paged + incremental tests green (identical output). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Summary
Device-agnostic execution + GGUF Q8_0 quantized serving + CPU decode/prefill performance, stacked on the unified continuous-batching engine (#1888).
Verified end-to-end on real SmolLM2-135M GGUF, llama.cpp parity preserved throughout (greedy next-token = 7042 " Paris" at every step).
Performance (SmolLM2-135M Q8_0, this AVX2 box)
What's here
GgufFile.TryReadQ8_0Raw->GgufModelSource.TryReadQ8_0->LlamaModelBuilder.LoadDense->DenseLayer.SetQuantizedWeightsQ8_0), running the block-Q8_0 GEMM (TensorsQ8BlockGemm) directly on int8. 3.76x less weight RAM. The quantized forward is gated on VNNI hardware (AvxVnni/Avx512) — on this AVX2 box it stays on fp32 (the int8 dot needsvpdpbusdto beat the tuned fp32 BLAS), auto-engaging on the CUDA+Linux serving target.PagedCachedMultiHeadAttention.MatVecMul, AVX2 FMA); bulk-read the paged KV history under one lock (PagedKVCache.ReadKeyValueRange) instead of a lock per position.model.To(DeviceInfo)/layer.To(...)placement,ConfigureModel(source, DeviceInfo)facade overload, GQA paged incremental wiring (InferenceOptimizerrecurses intoPreLNTransformerBlock).DEVHOST_PROFILE(steady-state prefill for dotnet-trace) andDEVHOST_DECODE(real incremental generation -> tok/s).Release gate
The pin is a local probe (
AiDotNet.Tensors 0.118.10-q8) — this consumer work depends on Tensorsfeat/q8-block-gemm(theQ8BlockGemmkernel + the float-RoPE/NumOps/mask-cache perf fixes, opened as a separate PR) plus #828 (already merged). CI will be red until that Tensors branch is released to nuget.org and the pin is bumped to the released version.🤖 Generated with Claude Code