PTXBench contains one GPU profiling service and two ways to run kernel agents: the multiturn model loop and the coding-agent gateway. Fixit SFT and KernelGen use the same multiturn runner and prompt registry for source and final evaluation. This repository gives the experiment process and a small set of runnable examples; it does not include paper outputs, plots, or the full prepared experiment matrix.
| Directory | Purpose |
|---|---|
fib-profile/ |
Five bundled H100 workload definitions and the GPU profiling service |
multiturn/ |
Model loop, CUDA/Triton evaluator tests, prompt registry, and runner |
ptxbench-eval/ |
Coding-agent gateway, launcher, agent images, and examples |
experiments/prepared/configs/ |
One multiturn config example |
experiments/fixit/ |
Failed-kernel repair, reasoning, SFT, serving, evaluation |
experiments/kernelgen/ |
Correct-kernel reasoning, SFT, serving, evaluation |
experiments/shared/ |
Run manifests, turn export, and shared training/serving code |
🚧 This repository is still under construction.
Use Python 3.12, Docker with NVIDIA GPU access, tmux, and an H100 host for the
bundled workloads. SFT runtime-SASS export also needs host nvcc,
cuobjdump, and TVM-FFI headers/libraries. The profiling service has its own
environment in the Docker image; install the host runner and coding-agent
launcher separately:
python3.12 -m venv .venv
source .venv/bin/activate
pip install -e '.[sft,dev]'
pip install -e ./ptxbench-eval
pip install -r multiturn/requirements.txtThe SFT extra installs Tinker Cookbook for training and checkpoint handling.
Training also needs a Tinker account. For evaluation alone, pip install -e .
and the multiturn requirements suffice. Model provider credentials are read from
environment variables; see .env.example.
Start the service using fib-profile/scripts/README.md. For the backward attention workloads, generate the two checksum-verified input blobs as described in fib-profile/dataset/README.md before starting the service. Check service health with:
curl -fsS http://localhost:10000/healthBuild the multiturn evaluator image:
docker build -f multiturn/docker/Dockerfile.eval -t ptxbench-multiturn-eval:latest .The runner expands a JSON config into trajectories and writes a plan.json,
trajectories, logs, and success artifacts into a new output root. The shared
prompt registry is multiturn/prompts/hub.json; assembled documents are under
multiturn/prompts/assembled. The same runner supports CUDA and Triton test
scripts under multiturn/tests/.
export GEMINI_API_KEY=... # for this example model
python multiturn/scripts/render.py --check
python multiturn/run_parallel_v2.py \
--config experiments/prepared/configs/example.json \
--definition gemm_n7168_k5120 \
--test-path multiturn/tests/cuda/gemm_n7168_k5120.py \
--model gemini-3.1-pro-preview \
--service-url http://localhost:10000 \
--gpu-arch hopper --without-local-gpu \
--max-parallel 1 --max-profiles 1 \
--image ptxbench-multiturn-eval:latest \
--output-root data/eval_runs/gemini-gemm-exampleResume an interrupted root with the same model, service, image, and output
options plus --resume, omitting config, definition, and test path. To run a
locally served Qwen model, set PTXBENCH_MODEL_HOST=host:port and use
--model Qwen3.6-27B; the endpoint must expose model ID
Qwen/Qwen3.6-27B via its OpenAI-compatible API.
ptxbench-eval uses the same prompt hub and profiling service. Its two H100
GEMM examples are under ptxbench-eval/examples/. The launchers build their
agent and gateway images, create a fresh experiment root, and run the gateway:
export GEMINI_API_KEY=...
bash ptxbench-eval/examples/gemini_gemm/run_experiment.shFor Codex, set up Codex authentication at ~/.codex/auth.json and run:
bash ptxbench-eval/examples/gpt56_gemm/run_experiment.shEach experiment.json supplies the three example prompt tags, model, workload,
and limits. Edit a copy for another coding-agent experiment. See
ptxbench-eval/WORKFLOW.md for prerequisites and
launcher options.
Fixit selects failed base-model kernels, asks Gemini to repair them, synthesizes reasoning, trains a LoRA adapter, serves it, and evaluates the result. KernelGen selects correct Gemini kernels, synthesizes reasoning with GLM-5.2, then uses the same training and evaluation path. Both workflows have ordered stages and machine-readable source/evaluation manifests:
bash scripts/reproduce_fixit.sh --check
bash scripts/reproduce_kernelgen.sh --check
bash scripts/reproduce_fixit.sh 00 # run one stage
bash scripts/reproduce_kernelgen.sh 00The example source manifests use four H100 d128 MHA workloads; evaluation also
includes the bundled GEMM. Adjust manifests and configs for a larger experiment.
Run stages in order using the details in Fixit
and KernelGen. Outputs default to data/;
set PTXBENCH_DATA_ROOT to persistent storage when running long experiments.
python multiturn/scripts/render.py --check
bash scripts/reproduce_fixit.sh --check
bash scripts/reproduce_kernelgen.sh --check
LITELLM_LOCAL_MODEL_COST_MAP=True pytest multiturn/tests experiments/testsStatic checks validate config and stage wiring. GPU profiling, hosted model calls, and Tinker training require their respective services and credentials.
