Skip to content
 
 

Latest commit

 

History

508 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Gufo: the Strix Halo inference engine

Gufo logo

Gufo is a vertical local inference engine specifically built and optimized for the AMD Strix Halo hardware: Ryzen AI MAX+ 395 systems with Radeon 8060S (gfx1151), up to 128 GiB of unified memory. This fork is specifically for adding Windows support for Gufo.

Special thanks to pixmaate for creating the initial windows port https://github.com/pixmaate/gufo and to the original Gufo team https://github.com/gufo-org/gufo

Models and benchmarks (benchmarks listed below were run on Linux)

All model documentation lives under docs/models:

Model Inference modes Hugging Face weights Benchmarks Quality
Qwen3.8 27B Q4/Q8, images, AR, DFlash2 Unsloth Q4_K_XL / Q8_K_XL · DFlash2 Q4_K_M Q4: 656.33 tok/s pp; up to 70.56 tok/s tg single user and 123.00 aggregated tok/s on 8 concurrent requests with DFlash2 · Benchmarks Quality
Qwen3.8 Flash-Next Q4, images, AR, MTP Unsloth Q4_K_XL · MTP Q8_0 1,628.52 tok/s pp; up to 59.41 tok/s tg single user and 157.22 aggregated tok/s on 8 concurrent requests with MTP · Benchmarks Quality
DeepSeek V4 Flash AR, DSpark antirez Flash 0731 IQ2XXS · DSpark 484.62 tok/s pp; up to 26.62 tok/s tg single user and 54.74 aggregated tok/s on 8 concurrent requests with DSpark · Benchmarks Quality
Qwen3-ASR 1.7B Speech recognition BF16 15.27× realtime · Benchmarks Quality
Qwen3-TTS 1.7B Speech synthesis and voice cloning BF16 CustomVoice / VoiceDesign / Base Up to 2.54× realtime; 201 ms to first audio (CustomVoice) · Benchmarks Quality
Qwen-Image-2.1 BF16 image generation and editing Complete pipeline In progress · Benchmarks Quality
MiniMax H3 BF16 text to video/audio FL2VA pipeline In progress · Benchmarks Quality

Peak measured workloads; text pp is autoregressive (AR), while tg uses the named speculative mode. Peaks include repetitive output; aggregate tg sums individual request decode rates. Qwen27B's single-user peak uses the short-prompt C1 workload. Audio excludes loading. Each model guide lists the required files and complete benchmark settings.

Philosophy

  • Contributions are welcome! We need the help of Strix Halo community to keep improving gufo!
  • We would like this to be the one-stop shop for Strix Halo Local AI enthusiasts: batteries included for text, audio, image, and video models.
  • Build and optimize specifically for the Strix Halo 128 GiB hardware. Smaller memory configurations should still work and preserve the speed benefits for models that can fit on memory.
  • Support only the best available models for their size that can run on this hardware: less code to maintain, more focused optimization and testing work.
  • Preserve quality when optimizing. Each model's quality report records independent numerical checks, execution consistency and unresolved gaps. Don't reuse kernels across different models to limit blast radius of a code change.
  • Treat concurrent requests, cancellation and conversation caching as first-class workloads.
  • Keep production dependencies small and development tools separate.

Windows quickstart

Run Qwen3.8-Flash-Next natively on Windows 11, Strix Halo / gfx1151, 128 GB RAM. Install the AMD graphics driver. For Flash-Next at 256K context, set 96 GB dedicated GPU memory in AMD Software → Performance → Tuning → Variable Graphics Memory. Setup checks the computer but does not change these settings.

  1. Install Git for Windows, then clone this fork:

    git clone --branch windows-port https://github.com/thomas9120/gufo.git
  2. Open the cloned gufo folder and double-click setup-windows.bat. Review the plan and enter y. Setup reuses existing tools, installs missing prerequisites, builds Gufo, and opens the GUI. Allow the installers to finish; the first setup needs substantial downloads, disk space, and compilation time. If interrupted, run the same file again to resume completed work.

  3. In the GUI, choose your model and optional MTP/projector files, adjust settings, click Save, then Start. For Flash-Next vision, use the compatible BF16 projector. See model download instructions if you do not already have the files. Connect your chat client to http://127.0.0.1:8080/v1 with the model name shown in the GUI (gufo by default).

For subsequent launches, double-click launch-gui.bat or a desktop shortcut to it. The GUI opens at http://127.0.0.1:8090; opening it does not load a model. Use Stop to unload Gufo or Exit to close both processes.

To update this checkout from GitHub and rebuild build\release\gufo.exe, use Exit in the GUI, then double-click update-windows.bat. It uses the current branch's configured remote and the existing build tools. See the update instructions for details.

See guided setup details for logs, preview mode, and custom dependency paths, or use the manual build instructions.

Linux quickstart

hf download unsloth/Qwen3.8-27B-GGUF \
  Qwen3.8-27B-UD-Q8_K_XL.gguf \
  --revision 4ca720788d1e01f1bff70c033e0d0028fd02e502 \
  --repo-type model \
  --local-dir models/Qwen3.8-27B-GGUF
hf download z-lab/Qwen3.8-27B-DFlash2-GGUF \
  Qwen3.8-27B-DFlash2-Q4_K_M.gguf \
  --revision 2d9571f8ce46e151f61c6499c99dee6079e1d610 \
  --repo-type model \
  --local-dir models/Qwen3.8-27B-DFlash2-GGUF
podman pull ghcr.io/gufo-org/toolboxes/gufo-runtime:latest
podman run --rm \
  --userns=keep-id:uid=1000,gid=1000 \
  --device /dev/kfd \
  --device /dev/dri \
  --group-add keep-groups \
  --ulimit memlock=-1 \
  -p 8080:8080 \
  -v ./models:/models:ro \
  ghcr.io/gufo-org/toolboxes/gufo-runtime:latest \
  gufo serve --host 0.0.0.0 --port 8080 llm \
  --model /models/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q8_K_XL.gguf \
  --speculative dflash2 \
  --dflash-model /models/Qwen3.8-27B-DFlash2-GGUF/Qwen3.8-27B-DFlash2-Q4_K_M.gguf

Rootless Podman needs crun for --group-add keep-groups. Your host user must have read/write access to /dev/kfd and /dev/dri/renderD*, usually through the render and video groups; log out and back in after changing membership. Container groups named video/render do not preserve host supplementary groups. Check id, ls -l /dev/kfd /dev/dri/renderD*, and podman info --format '{{.Host.OCIRuntime.Name}}' if ROCm reports no device. See Podman's rootless group-access guidance.

On Fedora or another SELinux-enforcing host, GPU enumeration can succeed while SELinux blocks mapping /dev/kfd, causing ROCr to report a misleading “Memory critical” error. Check the host audit log:

sudo ausearch -m avc -ts recent | grep -E '/dev/kfd|hsa_device_t'

If it shows a denied map for the container, Podman documents this fix:

sudo setsebool -P container_use_devices true

This persistently allows containers to access device labels for devices passed into them; it affects all containers on that host. Review that policy scope before enabling it. See Podman's device documentation and the SELinux container policy.

Then, from another terminal, ask it something through the OpenAI-compatible API:

curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen3.8-27B-UD-Q8_K_XL",
    "messages": [{"role": "user", "content": "Say something"}]
  }'

For Open WebUI, VS Code and OpenAI SDK clients, set the API base URL to http://localhost:8080/v1. Chat Completions supports text, images, tools and streaming; Responses supports text and streaming. See the API contract.

The text server uses the model's native context by default and generates until EOS or the context is full. --context N sets context capacity per session; --max-tokens N sets a default response limit that clients can override. Reasoning tokens count toward that response limit.

Build from source

Linux x86-64 on AMD Strix Halo (gfx1151) is the supported target. CMake owns one production configuration for both Nix and ordinary Linux builds. Tests, profilers, tuning executables and Python reference runners are not installed with the production package. No build.sh wrapper is needed.

Windows 11 x64 builds natively against AMD's TheRock ROCm; see Windows for the toolchain, a prerequisite check and the build script.

With Nix

nix build
./result/bin/gufo diagnose
./result/bin/gufo serve llm --model /path/to/model.gguf

flake.lock pins the dependencies. nix develop adds profiling, model-download and independent evaluation tools; these are not runtime requirements. Optional benchmark baselines are selected separately with nix shell .#ds4-reference, .#llama-cpp-reference or .#llama-cpp-mtp-reference; see benchmarking. See testing for the small hosted CI suite and explicit local quality checks.

Without Nix

Install a C++20 compiler, CMake 3.21+, Ninja, pkg-config and the following development libraries. The currently qualified toolchain is GCC 15.3 and ROCm 7.2.3. Attention and audio convolution kernels are compiled directly from HIP. Python, Triton/AOTriton, Composable Kernel and MIOpen are not production build or runtime requirements.

Dependency Used for
ROCm HIP compiler/runtime, hipBLAS, hipBLASLt, rocBLAS GPU execution and matrix multiplication
hipCUB, rocPRIM, rocWMMA headers Compiled GPU kernels
ICU, libcurl, OpenSSL, libpng, libjpeg Tokenization, HTTPS, hashing and images
FFmpeg and ffprobe Video/audio output; invoked as separate executables

Install ROCm using AMD's Linux instructions. Use the development packages for the libraries above. ROCm normally installs under /opt/rocm.

For example, on Debian/Ubuntu the ordinary system libraries are:

sudo apt install build-essential cmake ninja-build pkg-config \
  libicu-dev libcurl4-openssl-dev libssl-dev libpng-dev libjpeg-dev ffmpeg

# ROCm libraries from the table, named as AMD's repository ships them.
sudo apt install hipblas-dev hipblaslt-dev rocblas-dev \
  hipcub-dev rocprim-dev rocwmma-dev

cmake --preset release -DCMAKE_INSTALL_PREFIX="$HOME/.local"
cmake --build --preset release --parallel 4
./build/release/gufo diagnose
./build/release/gufo serve llm --model /path/to/model.gguf

Configuring fails at find_package(hipblas) when those ROCm packages are missing. Other distributions name them -devel instead of -dev. For nonstandard installations, pass ordinary CMake paths, for example cmake --preset release -DCMAKE_PREFIX_PATH="/opt/rocm". If compiler discovery picks a system Clang, also pass -DCMAKE_HIP_COMPILER=/opt/rocm/llvm/bin/clang++. Use cmake --install build/release to install Gufo, its runtime data and license notices. The GPU driver must allow your user to access /dev/kfd and /dev/dri; model weights are acquired separately.

The same source, compiler flags and install rules serve both builds. Nix pins the complete toolchain for reproducible comparisons; changing the compiler or math libraries requires the affected model's quality checks.

License

Gufo's original code is MIT licensed. Adapted code and dependencies retain their own notices in NOTICE, THIRD_PARTY_NOTICES.md and licenses/, installed under share/licenses/gufo. Model weights are not bundled and retain their publishers' terms.

Reference Projects

The initial design is informed by the following open source projects:

About

Windows port of Gufo's Strix Halo inference engine. Qwen Flash Next optimized. New GUI Launcher for Windows.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages