Prerequisites
Feature Description
I would like to request an option to use GPU-accelerated vision processing without keeping the complete multimodal projector permanently resident in VRAM.
Currently there are essentially two choices:
--mmproj-offload: fast GPU image encoding, but the projector permanently consumes a significant amount of VRAM.
--no-mmproj-offload: saves the VRAM, but image encoding runs on the CPU and can be considerably slower.
What I would like is a third option: keep the VRAM available for the LLM / KV cache most of the time, while still performing image encoding on the GPU.
There already seem to be several approaches demonstrating that this is possible:
ggml-org#20246
#140
https://github.com/valujin/beellama-kvarn
I also encountered a similar low-VRAM GPU vision approach while testing HyperQwen / vLLM.
Motivation
Currently, loading mmproj typically consumes around 1 GB of VRAM.
On VRAM-constrained GPUs, this is quite significant, especially when running large models with long context windows. That VRAM could otherwise be used for additional context, larger KV cache, draft models, or other features.
CPU-based image processing is an alternative, but in my experience it creates substantial overhead and is much slower.
For this reason, I currently often disable mmproj completely for Qwen and Gemma models: keeping it on the GPU costs too much VRAM, while processing images on the CPU is too expensive.
A low-memory GPU implementation would provide a useful middle ground:
-
fast GPU-based image processing
-
very low additional VRAM usage
-
more VRAM available for context and KV cache
-
no need to choose between fast vision processing and large context sizes
For users running modern multimodal models on GPUs with limited VRAM, this could make a substantial difference.
Possible Implementation
HyperQwen may provide a useful reference implementation. Its vLLM-based implementation seems to use a sliding-window approach for image processing, allowing the processing to remain GPU-accelerated while using only a very small amount of VRAM.
The related llama.cpp discussion also appears to cover approaches in this direction:
ggml-org#20246
I assume that a fully model-independent implementation may not be straightforward, since projector architectures differ between model families.
Still, even an implementation initially targeting a few common architectures such as Qwen and Gemma could already be very useful, and might provide a basis for a more generic solution later.
Prerequisites
Feature Description
I would like to request an option to use GPU-accelerated vision processing without keeping the complete multimodal projector permanently resident in VRAM.
Currently there are essentially two choices:
--mmproj-offload: fast GPU image encoding, but the projector permanently consumes a significant amount of VRAM.
--no-mmproj-offload: saves the VRAM, but image encoding runs on the CPU and can be considerably slower.
What I would like is a third option: keep the VRAM available for the LLM / KV cache most of the time, while still performing image encoding on the GPU.
There already seem to be several approaches demonstrating that this is possible:
ggml-org#20246
#140
https://github.com/valujin/beellama-kvarn
I also encountered a similar low-VRAM GPU vision approach while testing HyperQwen / vLLM.
Motivation
Currently, loading mmproj typically consumes around 1 GB of VRAM.
On VRAM-constrained GPUs, this is quite significant, especially when running large models with long context windows. That VRAM could otherwise be used for additional context, larger KV cache, draft models, or other features.
CPU-based image processing is an alternative, but in my experience it creates substantial overhead and is much slower.
For this reason, I currently often disable mmproj completely for Qwen and Gemma models: keeping it on the GPU costs too much VRAM, while processing images on the CPU is too expensive.
A low-memory GPU implementation would provide a useful middle ground:
fast GPU-based image processing
very low additional VRAM usage
more VRAM available for context and KV cache
no need to choose between fast vision processing and large context sizes
For users running modern multimodal models on GPUs with limited VRAM, this could make a substantial difference.
Possible Implementation
HyperQwen may provide a useful reference implementation. Its vLLM-based implementation seems to use a sliding-window approach for image processing, allowing the processing to remain GPU-accelerated while using only a very small amount of VRAM.
The related llama.cpp discussion also appears to cover approaches in this direction:
ggml-org#20246
I assume that a fully model-independent implementation may not be straightforward, since projector architectures differ between model families.
Still, even an implementation initially targeting a few common architectures such as Qwen and Gemma could already be very useful, and might provide a basis for a more generic solution later.