Conversation
Signed-off-by: Michel Belleau <michel.belleau@malaiwah.com>
|
Adversarial review follow-up is now in
Focused result on the pinned GG test image: 21 passed. |
Signed-off-by: Michel Belleau <michel.belleau@malaiwah.com>
|
Actual layer-3 producer pilot found a remaining gated-Qwen execution bug: the fused path serialized the correct 14,336-wide QG/K/V payload but |
Signed-off-by: Michel Belleau <michel.belleau@malaiwah.com>
|
The next actual pilot reached the extension fast path and exposed one more concrete contract: |
Purpose
Proof of concept for #297: allow one checkpoint to choose QKV topology per full-attention block.
split: existing q/k/v EXL3 objects, independent K/codebook/scale choices, preserving mixed precision and the serving consumer's three-launch route.fused_uniform: concatenate the BF16 q/k/v source weights in frozen q,k,v order, quantize that matrix jointly at one common K/codebook/scale, and serialize one qkv payload for a one-launch consumer.This is the practical hybrid requested by the downstream allocator: retain split widths where q/v sensitivity matters; use uniform fusion only where a new fused direct marginal shows no decision-relevant fidelity loss. A heterogeneous-width one-launch tuple remains future kernel/format work.
Implementation
Adds versioned
exl3_qkv_topology/1metadata and a deterministic converter plan:No load-time re-encoding and no duplicate split+fused payload for one block. Split remains the default.
Verification
The companion source verifier pins the patch/base/patched identities and mixed-schema invariants. Focused tests were executed in the immutable CUDA 13.2/r34 environment after compiling the pinned source extension:
Coverage includes mixed maps/runtime construction, BF16 concatenation/split bit identity, common-K/codebook/scale enforcement, unknown/duplicate declarations, and missing/duplicate payloads.
A consumer PoC is open at local-inference-lab/vllm#454. The ~88 us split versus ~36 us fused figure remains modeled; this PR enables the real block/end-logit/graph/TG experiment and makes no performance claim.
AI assistance was used to implement and review the PoC; the patch identity, extension build and focused tests were independently executed as described.