Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .claude/rules/project-layout.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@ The 4 files `kernel_compiler.py`, `runtime_compiler.py`, `toolchain.py`, `elf_pa
| What | Where |
| ---- | ----- |
| Runtime selection | `@scene_test(runtime="...")` on the SceneTestCase class |
| Per-case knobs (aicpu_thread_num, block_dim) | `CASES[*]["config"]` on the SceneTestCase class |
| Per-case knobs (aicpu_thread_num, runtime_env) | `CASES[*]["config"]` on the SceneTestCase class |
| Per-runtime build config | `src/{arch}/runtime/{runtime}/build_config.py` |
| Runtime build orchestration | `simpler_setup/runtime_builder.py` → `simpler_setup/runtime_compiler.py` → cmake |
| Pre-build all runtimes | `simpler_setup/build_runtimes.py` (invoked by `pip install .`) |
Expand Down
2 changes: 1 addition & 1 deletion .claude/skills/l0-swimlane/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -87,7 +87,7 @@ Verified examples (slots read from source):
### `--spmd-block-num N` — SPMD grid width

`block_num` is written into the synthesised slot-48 `LocalContext`. Default
is the case's `block_dim`; override only for a kernel that branches or
is 1; pass it to replay an SPMD cohort, or for a kernel that branches or
grid-strides on `block_num` (e.g. set the real hw width `24`). `block_idx`
is always synthesised to `0` (a representative block) and is **not** a flag —
it has no instruction-stream branches (doc §8).
Expand Down
4 changes: 2 additions & 2 deletions docs/aicore-kernel-programming.md
Original file line number Diff line number Diff line change
Expand Up @@ -88,10 +88,10 @@ that the runtime launches:

| Symbol | Meaning |
| ------ | ------- |
| `RUNTIME_CONFIG.block_dim` (Python `CallConfig.block_dim`) | Number of physical AICore blocks the runtime launches per dispatch. |
| `rt_available_cluster_count()` | Number of physical AICore blocks this run launches — the whole device; there is no per-call knob. |
| `get_block_num(args)` | Logical block count the kernel partitions work across. Currently always 1; multi-logical-block (`block_num > 1`) is not yet implemented. |

When you set `CallConfig.block_dim = 24` in Python and your kernel sees
When the device reports 24 clusters and your kernel sees
`get_block_num(args) == 1`, that is by design — every physical block
runs the same kernel and the kernel partitions work however it likes
using `get_block_idx()` against whatever it expects. Don't conflate
Expand Down
9 changes: 4 additions & 5 deletions docs/chip-level-arch.md
Original file line number Diff line number Diff line change
Expand Up @@ -106,7 +106,7 @@ DeviceRunner runner;
void *ptr = runner.allocate_tensor(bytes);
runner.copy_to_device(dev_ptr, host_ptr, bytes);
runner.set_executors(aicpu_binary, aicore_binary); // once, at init time
runner.run(runtime, config); // config carries block_dim, aicpu_thread_num, diagnostics
runner.run(runtime, config); // config carries aicpu_thread_num, diagnostics
runner.finalize();
```

Expand All @@ -124,7 +124,7 @@ simpler_init(ctx, device_id, // attach + binary takeove
size_t size = get_runtime_size();
register_callable(ctx, cid, callable); // one-time per callable
simpler_run(ctx, runtime, cid, args, config); // per-launch — no binaries; config
// carries block_dim, aicpu_thread_num,
// carries aicpu_thread_num,
// diagnostics + ring overrides
unregister_callable(ctx, cid);
finalize_device(ctx);
Expand All @@ -140,8 +140,7 @@ worker = ChipWorker()
worker.init(device_id=0, bins=bins) # bins = RuntimeBuilder(platform).get_binaries(...)

config = CallConfig()
# config.block_dim defaults to 0 = auto (DeviceRunner resolves to the max
# the AICore stream allows). Set explicitly to pin a smaller value.
# A run always takes the whole device; there is no per-call width knob.
config.aicpu_thread_num = 3
config.enable_pmu = 0
worker.run(callable, args, config)
Expand Down Expand Up @@ -204,7 +203,7 @@ binding — they must be called from the same thread that called `init()`.
### 3. Execution Phase

```text
worker.run(callable, args, CallConfig(block_dim, aicpu_thread_num))
worker.run(callable, args, CallConfig(aicpu_thread_num))
│
└─→ run_runtime(ctx, runtime, callable, args, ...)
│
Expand Down
19 changes: 10 additions & 9 deletions docs/dfx/l0-swimlane-profiling.md
Original file line number Diff line number Diff line change
Expand Up @@ -135,7 +135,7 @@ reading the kernel source — names are not in the dump (only kind / shape
/ value), so cross-reference the kernel's `args:` header for those:

```text
[l0_swimlane] func_id=0 task=0x... mix=[0, 1, 2] mode=mix block_dim=3
[l0_swimlane] func_id=0 task=0x... mix=[0, 1, 2] mode=mix block_num=3
members=[MATMUL(aic,func 0), ADD(aiv,func 1), MUL(aiv,func 2)]
[l0_swimlane] arg slots (override with --set-arg SLOT=VALUE):
slot 0 tensor FLOAT32 [16384]
Expand Down Expand Up @@ -231,7 +231,7 @@ Omitting `--case` auto-pins the **first** `CASES[*]` that lists your
no "run every case, reconstruct from the newest dump dir" ambiguity). Pass
`--case` explicitly when that first case is not the smallest — a full-size
production case's shapes overflow the camodel replay (§3.4). The synthesized
slot-48 `block_num` is taken from the **selected** case's `block_dim`.
slot-48 `block_num` comes from `--spmd-block-num` (default 1); an SPMD cohort sizes itself on device, so it is not readable from the test file.

### 3.6 Reusing a dump across kernels

Expand Down Expand Up @@ -267,7 +267,7 @@ arch-precheck (the case must declare the `--platform` you pass — §3.1).
| Single AIV | `vector_example` | `--func-id 0` | `kernel_add`, dispatched `rt_submit_aiv_task(0)` (vec only) |
| Mix 2 AIV (per-lane) | `mixed_example` | `--func-id 3,4` | ADD_STD@AIV0 + MUL_STD@AIV1 (`get_subblockid` routing) |
| Mix 3-way 1C2V | `mixed_example` | `--func-id 0,1,2` | MATMUL@AIC + ADD@AIV0 + MUL@AIV1 |
| SPMD single-source | `spmd_multiblock_aiv` | `--func-id 0` | single AIV reading `get_block_idx` (`block_dim=24`; replay traces block 0) |
| SPMD single-source | `spmd_multiblock_aiv` | `--func-id 0` | single AIV reading `get_block_idx` (pass `--spmd-block-num`; replay traces block 0) |
| SPMD mix, 2 AIV share a source | `spmd_multiblock_mix` | `--func-id 0,1,2` | func 1 & 2 are distinct ids but **both `kernel_spmd_mix.cpp`** → the 2 AIV collapse to one (both lanes run it). Routes by `get_sub_block_id` (slot 49) → in replay both lanes read `sub_block_id=0`; AIV0/AIV1 differ only by write offset, so the pipeline stays representative. (The same-source collapse also covers the duplicate-func_id `[0,1,1]` shape an SPMD mix produces when `aiv0 = aiv1`.) |
| Paged-attn, loop = scalar | `paged_attention_unroll` | `--func-id 0 --set-arg 4=4` | QK stage; `n_blocks` scalar (slot 4) → shrink to 4 ([§7.2](#72---set-arg-floor-for-a-loop-count-without-distortion)) |
| Paged-attn, loop = control tensor | `batch_paged_attention` | `--func-id 1 --set-arg 1=512 --case CaseSmall1` | SF reads `context_lens` (**slot 1**) content (`aiv_softmax_prepare.cpp`); `--set-arg 1=512` fills it uniformly → shrinks the derived per-batch block count |
Expand Down Expand Up @@ -441,15 +441,16 @@ SPMD kernels read an execution context the orchestration builds per
dispatch — `LocalContext{block_idx, block_num}` at args slot 48 and
`GlobalContext{sub_block_id}` at slot 49. The isolated replay has no
orchestration, so `replay_host.cpp` **synthesizes** it: one
`LocalContext{block_idx=0, block_num=block_dim}` + `GlobalContext`
`LocalContext{block_idx=0, block_num=--spmd-block-num}` + `GlobalContext`
pointed at slots 48/49. This is harmless for positional kernels (they
ignore 48/49) and required for SPMD kernels that read `get_block_idx` /
`get_block_num` (which would otherwise dereference null). `block_idx=0`
traces a representative block; `block_num = block_dim` (the **selected**
case's grid width — the `--case` case, else the auto-pinned first-platform
case) keeps steady-state branches (`block_idx+1 < block_num`) on their
normal path — see [§8](#8-fidelity-rules). `--spmd-block-num` overrides
`block_num`.
traces a representative block; `block_num` (from `--spmd-block-num`, default

1) keeps steady-state branches (`block_idx+1 < block_num`) on their normal
path — see [§8](#8-fidelity-rules). An SPMD cohort sizes itself from
`rt_available_cluster_count()` on device, so the width is not readable from
the test file and must be passed.
Note the per-AIV-lane routing for a mix uses the hardware
`get_subblockid()` (§5.4), not the synthesized slot-49 value.

Expand Down
2 changes: 1 addition & 1 deletion docs/dynamic-linking.md
Original file line number Diff line number Diff line change
Expand Up @@ -299,7 +299,7 @@ ChipWorker.run(handle, args, config) # public wrapper path
simpler_run(ctx, buf, internal callable entry, args, config)
new (buf) Runtime()
DeviceRunner::bind_callable_to_runtime(r, cid, api, args, rings) # replay + per-run bind
DeviceRunner::run(r, config) # applies config, resolves block_dim
DeviceRunner::run(r, config) # applies config; width already resolved pre-bind
clear_cpu_sim_shared_storage()
ensure_binaries_loaded() dlopen aicpu/aicore SOs once
launch AICPU + AICore threads
Expand Down
6 changes: 2 additions & 4 deletions docs/getting-started.md
Original file line number Diff line number Diff line change
Expand Up @@ -184,9 +184,8 @@ worker.init()
# Register the ChipCallable to obtain an opaque callable handle.
handle = worker.register(chip_callable)

# Execute the registered callable on device. Omitting block_dim uses the
# default 0 = auto, which DeviceRunner resolves to the max the AICore
# stream allows. Pass block_dim=<n> to pin a smaller value.
# Execute the registered callable on device. A run always takes the whole
# device; orchestration reads the resulting width via rt_available_cluster_count().
worker.run(handle, orch_args)

# Cleanup
Expand Down Expand Up @@ -216,7 +215,6 @@ Runtime behavior is configured via `kernel_config.py` in each example:
RUNTIME_CONFIG = {
"runtime": "host_build_graph", # Runtime to use
"aicpu_thread_num": 3, # Number of AICPU scheduler threads
"block_dim": 3, # Number of AICore blocks (1 block = 1 AIC + 2 AIV)
}
```

Expand Down
5 changes: 3 additions & 2 deletions docs/remote-l3-worker-design/protocol.md
Original file line number Diff line number Diff line change
Expand Up @@ -69,10 +69,11 @@ Encoding rules:
- Reserved fields must be written as zero and rejected when non-zero unless a
later protocol version assigns them.

`CallConfigWire v1` encodes the current `CallConfig` fields explicitly:
`CallConfigWire v2` encodes the current `CallConfig` fields explicitly (v1
carried a leading `block_dim: int32`, dropped when a run became
whole-device):

```text
block_dim: int32
aicpu_thread_num: int32
enable_l2_swimlane: int32
enable_dump_args: int32
Expand Down
5 changes: 2 additions & 3 deletions docs/task-flow.md
Original file line number Diff line number Diff line change
Expand Up @@ -35,7 +35,7 @@ Every task flowing through any level carries exactly three pieces of data:
| ------ | ---- | ---------- |
| `CallableHandle` / `CallableIdentity` | hash digest + kind + namespace | What the target worker should execute; targets resolve the digest to a local slot |
| `TaskArgs` | user builder class | Tensors + scalars + per-tensor tags (IN/OUT/INOUT/etc.) |
| `CallConfig` | small POD | Execution knobs (block_dim, aicpu_thread_num, profiling/dump/PMU flags, …) |
| `CallConfig` | small POD | Execution knobs (aicpu_thread_num, profiling/dump/PMU flags, …) |

Everything else in the engine is either plumbing (slots, ring, tensormap,
scheduler) or target-local executable state resolved from the callable digest.
Expand Down Expand Up @@ -204,7 +204,6 @@ View does **not** own memory. Valid for the duration of a single

```cpp
struct CallConfig {
int32_t block_dim = 0; // 0 = auto (DeviceRunner resolves to stream max at run() time)
int32_t aicpu_thread_num = 3;
int32_t enable_l2_swimlane = 0; // perf_level 0–4 (0=off, 4=full)
int32_t enable_dump_args = 0;
Expand Down Expand Up @@ -537,7 +536,7 @@ w3 = Worker(level=3, child_mode=PROCESS)
w3.add_worker(NEXT_LEVEL, chip_worker_0)
w3.init() # fork chip_0 here

w3.run(my_orch, args, CallConfig(block_dim=3))
w3.run(my_orch, args, CallConfig(aicpu_thread_num=3))
```

Step-by-step (one chip worker):
Expand Down
2 changes: 1 addition & 1 deletion docs/testing.md
Original file line number Diff line number Diff line change
Expand Up @@ -596,7 +596,7 @@ class TestMyKernel(SceneTestCase):
{
"name": "default",
"platforms": ["a2a3sim", "a2a3"],
"config": {"aicpu_thread_num": 4, "block_dim": 3},
"config": {"aicpu_thread_num": 4},
"params": {},
},
]
Expand Down
6 changes: 3 additions & 3 deletions docs/troubleshooting/sim-oversubscription-hang.md
Original file line number Diff line number Diff line change
Expand Up @@ -43,8 +43,8 @@ easy to trip on any box running concurrent sim tests.
`--max-parallel` is the number of scene-test **cases** the scheduler runs
concurrently (`auto` = `min(devices, nproc)`). Each case forks one chip
subprocess per device, and each chip spawns one host thread per simulated
AICore (`block_dim` threads) plus AICPU + scheduler threads. So N concurrent
cases on an M-vCPU runner put roughly `N × devices × block_dim` runnable threads
AICore (`SIM_AUTO_BLOCKDIM` threads) plus AICPU + scheduler threads. So N concurrent
cases on an M-vCPU runner put roughly `N × devices × SIM_AUTO_BLOCKDIM` runnable threads
on M cores — easily 10–100× oversubscription.

Under that pressure two HW-correct patterns misbehave:
Expand Down Expand Up @@ -76,7 +76,7 @@ the sim host-thread model.
- **Throttle parallelism**: `pytest ... --max-parallel 2` (or lower) on a
CPU-constrained runner. This is the flag's intended use — it shrinks the
oversubscription without changing `--device`.
- **Shrink the per-case footprint**: a smaller `block_dim` in the case config
- **Shrink the per-case footprint**: `SIM_AUTO_BLOCKDIM` in the sim device runner
means fewer AICore threads per chip.

## Fix
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -52,41 +52,41 @@ class TestBenchmarkBgemm(SceneTestCase):
{
"name": "Case0",
"platforms": ["a2a3sim", "a2a3"],
"config": {"aicpu_thread_num": 4, "block_dim": 24},
"config": {"aicpu_thread_num": 4},
"params": {"matmul_add_task_num": 500, "incore_data_size": 128, "incore_loop": 4, "grid_k": 2},
},
{
"name": "Case1",
"manual": True,
"platforms": ["a2a3sim", "a2a3"],
"config": {"aicpu_thread_num": 4, "block_dim": 24},
"config": {"aicpu_thread_num": 4},
"params": {"matmul_add_task_num": 64, "incore_data_size": 128, "incore_loop": 4, "grid_k": 2},
},
{
"name": "Case2",
"manual": True,
"platforms": ["a2a3sim", "a2a3"],
"config": {"aicpu_thread_num": 4, "block_dim": 24},
"config": {"aicpu_thread_num": 4},
"params": {"matmul_add_task_num": 256, "incore_data_size": 128, "incore_loop": 4, "grid_k": 2},
},
{
"name": "Case3",
"manual": True,
"platforms": ["a2a3sim", "a2a3"],
"config": {"aicpu_thread_num": 4, "block_dim": 24},
"config": {"aicpu_thread_num": 4},
"params": {"matmul_add_task_num": 64, "incore_data_size": 128, "incore_loop": 16, "grid_k": 2},
},
{
"name": "Case4",
"manual": True,
"platforms": ["a2a3sim", "a2a3"],
"config": {"aicpu_thread_num": 4, "block_dim": 24},
"config": {"aicpu_thread_num": 4},
"params": {"matmul_add_task_num": 64, "incore_data_size": 128, "incore_loop": 4, "grid_k": 4},
},
{
"name": "Bgemm64",
"platforms": ["a2a3sim", "a2a3"],
"config": {"aicpu_thread_num": 4, "block_dim": 3},
"config": {"aicpu_thread_num": 4},
"params": {"matmul_add_task_num": 32, "incore_data_size": 64, "incore_loop": 1, "grid_k": 4},
},
]
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -61,7 +61,7 @@ class TestMergePipelineBarrier(SceneTestCase):
# per-task via launch_spec.set_block_num in merge_orch.cpp).
"name": "merge",
"platforms": ["a2a3"],
"config": {"aicpu_thread_num": 4, "block_dim": 24},
"config": {"aicpu_thread_num": 4},
"params": {},
},
]
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -64,7 +64,7 @@ class TestPagedAttention(SceneTestCase):
{
"name": "Case1",
"platforms": ["a2a3"],
"config": {"aicpu_thread_num": 4, "block_dim": 24},
"config": {"aicpu_thread_num": 4},
"params": {
"batch": 256,
"num_heads": 16,
Expand All @@ -79,7 +79,7 @@ class TestPagedAttention(SceneTestCase):
{
"name": "Case2",
"platforms": ["a2a3"],
"config": {"aicpu_thread_num": 4, "block_dim": 24},
"config": {"aicpu_thread_num": 4},
"manual": True,
"params": {
"batch": 64,
Expand All @@ -95,7 +95,7 @@ class TestPagedAttention(SceneTestCase):
{
"name": "Case3",
"platforms": ["a2a3"],
"config": {"aicpu_thread_num": 4, "block_dim": 24},
"config": {"aicpu_thread_num": 4},
"manual": True,
"params": {
"batch": 64,
Expand All @@ -111,7 +111,7 @@ class TestPagedAttention(SceneTestCase):
{
"name": "CaseSmall1",
"platforms": ["a2a3sim", "a2a3"],
"config": {"aicpu_thread_num": 4, "block_dim": 9},
"config": {"aicpu_thread_num": 4},
"params": {
"batch": 1,
"num_heads": 16,
Expand All @@ -126,7 +126,7 @@ class TestPagedAttention(SceneTestCase):
{
"name": "CaseSmall2",
"platforms": ["a2a3sim", "a2a3"],
"config": {"aicpu_thread_num": 4, "block_dim": 24},
"config": {"aicpu_thread_num": 4},
"manual": True,
"params": {
"batch": 1,
Expand All @@ -142,7 +142,7 @@ class TestPagedAttention(SceneTestCase):
{
"name": "CaseVarSeq2",
"platforms": ["a2a3sim", "a2a3"],
"config": {"aicpu_thread_num": 4, "block_dim": 24},
"config": {"aicpu_thread_num": 4},
"manual": True,
"params": {
"batch": 2,
Expand All @@ -159,7 +159,7 @@ class TestPagedAttention(SceneTestCase):
{
"name": "CaseVarSeq4",
"platforms": ["a2a3sim", "a2a3"],
"config": {"aicpu_thread_num": 4, "block_dim": 24},
"config": {"aicpu_thread_num": 4},
"manual": True,
"params": {
"batch": 4,
Expand Down
Loading
Loading