Gufo is a vertical local inference engine specifically built and optimized for the AMD Strix Halo hardware:
Ryzen AI MAX+ 395 systems with Radeon 8060S (gfx1151), up to 128 GiB of unified memory. This fork is specifically for adding Windows support for Gufo.
Special thanks to pixmaate for creating the initial windows port https://github.com/pixmaate/gufo and to the original Gufo team https://github.com/gufo-org/gufo
All model documentation lives under docs/models:
| Model | Inference modes | Hugging Face weights | Benchmarks | Quality |
|---|---|---|---|---|
| Qwen3.8 27B | Q4/Q8, images, AR, DFlash2 | Unsloth Q4_K_XL / Q8_K_XL · DFlash2 Q4_K_M | Q4: 656.33 tok/s pp; up to 70.56 tok/s tg single user and 123.00 aggregated tok/s on 8 concurrent requests with DFlash2 · Benchmarks | Quality |
| Qwen3.8 Flash-Next | Q4, images, AR, MTP | Unsloth Q4_K_XL · MTP Q8_0 | 1,628.52 tok/s pp; up to 59.41 tok/s tg single user and 157.22 aggregated tok/s on 8 concurrent requests with MTP · Benchmarks | Quality |
| DeepSeek V4 Flash | AR, DSpark | antirez Flash 0731 IQ2XXS · DSpark | 484.62 tok/s pp; up to 26.62 tok/s tg single user and 54.74 aggregated tok/s on 8 concurrent requests with DSpark · Benchmarks | Quality |
| Qwen3-ASR 1.7B | Speech recognition | BF16 | 15.27× realtime · Benchmarks | Quality |
| Qwen3-TTS 1.7B | Speech synthesis and voice cloning | BF16 CustomVoice / VoiceDesign / Base | Up to 2.54× realtime; 201 ms to first audio (CustomVoice) · Benchmarks | Quality |
| Qwen-Image-2.1 | BF16 image generation and editing | Complete pipeline | In progress · Benchmarks | Quality |
| MiniMax H3 | BF16 text to video/audio | FL2VA pipeline | In progress · Benchmarks | Quality |
Peak measured workloads; text pp is autoregressive (AR), while tg uses the named speculative mode. Peaks include repetitive output; aggregate tg sums individual request decode rates. Qwen27B's single-user peak uses the short-prompt C1 workload. Audio excludes loading. Each model guide lists the required files and complete benchmark settings.
- Contributions are welcome! We need the help of Strix Halo community to keep improving gufo!
- We would like this to be the one-stop shop for Strix Halo Local AI enthusiasts: batteries included for text, audio, image, and video models.
- Build and optimize specifically for the Strix Halo 128 GiB hardware. Smaller memory configurations should still work and preserve the speed benefits for models that can fit on memory.
- Support only the best available models for their size that can run on this hardware: less code to maintain, more focused optimization and testing work.
- Preserve quality when optimizing. Each model's quality report records independent numerical checks, execution consistency and unresolved gaps. Don't reuse kernels across different models to limit blast radius of a code change.
- Treat concurrent requests, cancellation and conversation caching as first-class workloads.
- Keep production dependencies small and development tools separate.
Run Qwen3.8-Flash-Next natively on Windows 11, Strix Halo / gfx1151, 128 GB RAM. Install the AMD graphics driver. For Flash-Next at 256K context, set 96 GB dedicated GPU memory in AMD Software → Performance → Tuning → Variable Graphics Memory. Setup checks the computer but does not change these settings.
-
Install Git for Windows, then clone this fork:
git clone --branch windows-port https://github.com/thomas9120/gufo.git
-
Open the cloned
gufofolder and double-click setup-windows.bat. Review the plan and entery. Setup reuses existing tools, installs missing prerequisites, builds Gufo, and opens the GUI. Allow the installers to finish; the first setup needs substantial downloads, disk space, and compilation time. If interrupted, run the same file again to resume completed work. -
In the GUI, choose your model and optional MTP/projector files, adjust settings, click Save, then Start. For Flash-Next vision, use the compatible BF16 projector. See model download instructions if you do not already have the files. Connect your chat client to
http://127.0.0.1:8080/v1with the model name shown in the GUI (gufoby default).
For subsequent launches, double-click launch-gui.bat or a
desktop shortcut to it. The GUI opens at
http://127.0.0.1:8090; opening it does not load a model. Use Stop to unload
Gufo or Exit to close both processes.
To update this checkout from GitHub and rebuild build\release\gufo.exe, use
Exit in the GUI, then double-click update-windows.bat.
It uses the current branch's configured remote and the existing build tools. See the
update instructions for details.
See guided setup details for logs, preview mode, and custom dependency paths, or use the manual build instructions.
hf download unsloth/Qwen3.8-27B-GGUF \
Qwen3.8-27B-UD-Q8_K_XL.gguf \
--revision 4ca720788d1e01f1bff70c033e0d0028fd02e502 \
--repo-type model \
--local-dir models/Qwen3.8-27B-GGUF
hf download z-lab/Qwen3.8-27B-DFlash2-GGUF \
Qwen3.8-27B-DFlash2-Q4_K_M.gguf \
--revision 2d9571f8ce46e151f61c6499c99dee6079e1d610 \
--repo-type model \
--local-dir models/Qwen3.8-27B-DFlash2-GGUF
podman pull ghcr.io/gufo-org/toolboxes/gufo-runtime:latest
podman run --rm \
--userns=keep-id:uid=1000,gid=1000 \
--device /dev/kfd \
--device /dev/dri \
--group-add keep-groups \
--ulimit memlock=-1 \
-p 8080:8080 \
-v ./models:/models:ro \
ghcr.io/gufo-org/toolboxes/gufo-runtime:latest \
gufo serve --host 0.0.0.0 --port 8080 llm \
--model /models/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q8_K_XL.gguf \
--speculative dflash2 \
--dflash-model /models/Qwen3.8-27B-DFlash2-GGUF/Qwen3.8-27B-DFlash2-Q4_K_M.ggufRootless Podman needs crun for --group-add keep-groups. Your host user must
have read/write access to /dev/kfd and /dev/dri/renderD*, usually through the
render and video groups; log out and back in after changing membership.
Container groups named video/render do not preserve host supplementary
groups. Check id, ls -l /dev/kfd /dev/dri/renderD*, and
podman info --format '{{.Host.OCIRuntime.Name}}' if ROCm reports no device.
See Podman's rootless group-access guidance.
On Fedora or another SELinux-enforcing host, GPU enumeration can succeed while
SELinux blocks mapping /dev/kfd, causing ROCr to report a misleading
“Memory critical” error. Check the host audit log:
sudo ausearch -m avc -ts recent | grep -E '/dev/kfd|hsa_device_t'If it shows a denied map for the container, Podman documents this fix:
sudo setsebool -P container_use_devices trueThis persistently allows containers to access device labels for devices passed into them; it affects all containers on that host. Review that policy scope before enabling it. See Podman's device documentation and the SELinux container policy.
Then, from another terminal, ask it something through the OpenAI-compatible API:
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen3.8-27B-UD-Q8_K_XL",
"messages": [{"role": "user", "content": "Say something"}]
}'For Open WebUI, VS Code and OpenAI SDK clients, set the API base URL to
http://localhost:8080/v1. Chat Completions supports text, images, tools and
streaming; Responses supports text and streaming. See the API contract.
The text server uses the model's native context by default and generates until
EOS or the context is full. --context N sets context capacity per session;
--max-tokens N sets a default response limit that clients can override.
Reasoning tokens count toward that response limit.
Linux x86-64 on AMD Strix Halo (gfx1151) is the supported target. CMake owns
one production configuration for both Nix and ordinary Linux builds. Tests,
profilers, tuning executables and Python reference runners are not installed
with the production package. No build.sh wrapper is needed.
Windows 11 x64 builds natively against AMD's TheRock ROCm; see Windows for the toolchain, a prerequisite check and the build script.
nix build
./result/bin/gufo diagnose
./result/bin/gufo serve llm --model /path/to/model.ggufflake.lock pins the dependencies. nix develop adds profiling,
model-download and independent evaluation tools; these are not runtime
requirements. Optional benchmark baselines are selected separately with
nix shell .#ds4-reference, .#llama-cpp-reference or
.#llama-cpp-mtp-reference; see benchmarking.
See testing for the small hosted CI suite and
explicit local quality checks.
Install a C++20 compiler, CMake 3.21+, Ninja, pkg-config and the following development libraries. The currently qualified toolchain is GCC 15.3 and ROCm 7.2.3. Attention and audio convolution kernels are compiled directly from HIP. Python, Triton/AOTriton, Composable Kernel and MIOpen are not production build or runtime requirements.
| Dependency | Used for |
|---|---|
| ROCm HIP compiler/runtime, hipBLAS, hipBLASLt, rocBLAS | GPU execution and matrix multiplication |
| hipCUB, rocPRIM, rocWMMA headers | Compiled GPU kernels |
| ICU, libcurl, OpenSSL, libpng, libjpeg | Tokenization, HTTPS, hashing and images |
| FFmpeg and ffprobe | Video/audio output; invoked as separate executables |
Install ROCm using AMD's Linux instructions.
Use the development packages for the libraries above. ROCm normally installs
under /opt/rocm.
For example, on Debian/Ubuntu the ordinary system libraries are:
sudo apt install build-essential cmake ninja-build pkg-config \
libicu-dev libcurl4-openssl-dev libssl-dev libpng-dev libjpeg-dev ffmpeg
# ROCm libraries from the table, named as AMD's repository ships them.
sudo apt install hipblas-dev hipblaslt-dev rocblas-dev \
hipcub-dev rocprim-dev rocwmma-dev
cmake --preset release -DCMAKE_INSTALL_PREFIX="$HOME/.local"
cmake --build --preset release --parallel 4
./build/release/gufo diagnose
./build/release/gufo serve llm --model /path/to/model.ggufConfiguring fails at find_package(hipblas) when those ROCm packages are
missing. Other distributions name them -devel instead of -dev.
For nonstandard installations, pass ordinary CMake paths, for example
cmake --preset release -DCMAKE_PREFIX_PATH="/opt/rocm".
If compiler discovery picks a system Clang, also pass
-DCMAKE_HIP_COMPILER=/opt/rocm/llvm/bin/clang++.
Use cmake --install build/release to install Gufo,
its runtime data and license notices. The GPU driver must allow your user to
access /dev/kfd and /dev/dri; model weights are acquired separately.
The same source, compiler flags and install rules serve both builds. Nix pins the complete toolchain for reproducible comparisons; changing the compiler or math libraries requires the affected model's quality checks.
Gufo's original code is MIT licensed. Adapted code and dependencies
retain their own notices in NOTICE, THIRD_PARTY_NOTICES.md
and licenses/, installed under share/licenses/gufo. Model weights are not
bundled and retain their publishers' terms.
The initial design is informed by the following open source projects:
- llama.cpp for compact model serving, GGUF, and CPU/GPU correctness paths.
- LaurentZuijdwijk/llama.cpp, Nathanw1014/strix-halo-llamacpp, and gaetan-puleo/llama-cpp-strix-halo for Strix Halo optimization inspiration.
- vLLM for continuous batching and paged request scheduling.
- hipEngine for torch-free HIP execution, and native speculative-cycle work.
- ds4 for DeepSeek V4 Flash, MoE scheduling, and DSpark.
- audio.cpp for audio models for tts and asr tasks.
