Skip to content

PoC: per-layer split or fused-uniform QKV topology - #299

Open
malaiwah wants to merge 4 commits into
turboderp-org:masterfrom
malaiwah:poc/per-layer-qkv-topology
Open

malaiwah wants to merge 4 commits into
turboderp-org:masterfrom
malaiwah:poc/per-layer-qkv-topology

Conversation

@malaiwah

Copy link
Copy Markdown
Contributor

Purpose

Proof of concept for #297: allow one checkpoint to choose QKV topology per full-attention block.

  • split: existing q/k/v EXL3 objects, independent K/codebook/scale choices, preserving mixed precision and the serving consumer's three-launch route.
  • fused_uniform: concatenate the BF16 q/k/v source weights in frozen q,k,v order, quantize that matrix jointly at one common K/codebook/scale, and serialize one qkv payload for a one-launch consumer.

This is the practical hybrid requested by the downstream allocator: retain split widths where q/v sensitivity matters; use uniform fusion only where a new fused direct marginal shows no decision-relevant fidelity loss. A heterogeneous-width one-launch tuple remains future kernel/format work.

Implementation

Adds versioned exl3_qkv_topology/1 metadata and a deterministic converter plan:

  • complete split-default rows with per-layer fused overrides;
  • direct BF16 q/k/v loading and exact q,k,v-order concatenation;
  • one joint quantization call retaining the input calibration qmap;
  • mixed split/fused Attention construction and runtime routing;
  • config, safetensors/index and tensor-storage logical reconstruction metadata;
  • fail-closed common-K/codebook/scale, unknown/duplicate declaration, missing/duplicate payload and component checks;
  • fused TP explicitly rejected by the PoC rather than silently mis-sharded.

No load-time re-encoding and no duplicate split+fused payload for one block. Split remains the default.

Verification

The companion source verifier pins the patch/base/patched identities and mixed-schema invariants. Focused tests were executed in the immutable CUDA 13.2/r34 environment after compiling the pinned source extension:

6 passed, 1 warning in 1.89s

Coverage includes mixed maps/runtime construction, BF16 concatenation/split bit identity, common-K/codebook/scale enforcement, unknown/duplicate declarations, and missing/duplicate payloads.

A consumer PoC is open at local-inference-lab/vllm#454. The ~88 us split versus ~36 us fused figure remains modeled; this PR enables the real block/end-logit/graph/TG experiment and makes no performance claim.

AI assistance was used to implement and review the PoC; the patch identity, extension build and focused tests were independently executed as described.

Signed-off-by: Michel Belleau <michel.belleau@malaiwah.com>
@malaiwah

Copy link
Copy Markdown
Contributor Author

Adversarial review follow-up is now in 3777614a5af305d8e2db4bab87f940983e9c71e0.

  • No-plan conversion now short-circuits topology handling and preserves existing behavior, including ordinary 3inst conversions.
  • Opted-in topology handles Qwen3.5/3.8 interleaved Q/G with the real doubled q projection width ([12288,1024,1024] for the target).
  • Schema rows/completeness are scoped to the primary text decoder; vision and MTP stay on their established payloads.
  • The producer's opted-in schema domain now matches the consumer: K3..K8 and mcg/mul1, with lexical full-path ordering.
  • Added realistic interleaved, side-model exclusion, no-plan, and domain-boundary tests.

Focused result on the pinned GG test image: 21 passed.

Signed-off-by: Michel Belleau <michel.belleau@malaiwah.com>
@malaiwah

Copy link
Copy Markdown
Contributor Author

Actual layer-3 producer pilot found a remaining gated-Qwen execution bug: the fused path serialized the correct 14,336-wide QG/K/V payload but Attention.project_qkv still split only 6,144 Q values. 0f92b6f4 now splits the complete 12,288-wide interleaved Q/G component, deinterleaves it through the established fast/fallback paths, and returns the 24×256 query plus 6,144 gate values. Added a forward regression over the real 24/4/256 geometry. Focused producer tests: 22 passed; combined producer+GDN source tests: 26 passed.

Signed-off-by: Michel Belleau <michel.belleau@malaiwah.com>
@malaiwah

Copy link
Copy Markdown
Contributor Author

The next actual pilot reached the extension fast path and exposed one more concrete contract: torch.split returns a non-contiguous Q/G view, while ext.deinterleave_qg requires contiguous input. dce9cfa materializes only that fused Q/G slice before the existing deinterleave kernel and adds a half-precision fast-path regression that asserts contiguity. Producer tests: 23 passed; combined producer+GDN tests: 27 passed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant