Skip to content

Feature Request: mmproj GPU processing while offloading weights #179

Description

@katakombi

Prerequisites

  • I am running the latest code. Mention the version if possible as well.
  • I carefully followed the README.md.
  • I searched using keywords relevant to my issue to make sure that I am creating a new issue that is not already open (or closed).
  • I reviewed the Discussions, and have a new and useful enhancement to share.

Feature Description

I would like to request an option to use GPU-accelerated vision processing without keeping the complete multimodal projector permanently resident in VRAM.

Currently there are essentially two choices:

--mmproj-offload: fast GPU image encoding, but the projector permanently consumes a significant amount of VRAM.

--no-mmproj-offload: saves the VRAM, but image encoding runs on the CPU and can be considerably slower.

What I would like is a third option: keep the VRAM available for the LLM / KV cache most of the time, while still performing image encoding on the GPU.

There already seem to be several approaches demonstrating that this is possible:

ggml-org#20246

#140

https://github.com/valujin/beellama-kvarn

I also encountered a similar low-VRAM GPU vision approach while testing HyperQwen / vLLM.

Motivation

Currently, loading mmproj typically consumes around 1 GB of VRAM.

On VRAM-constrained GPUs, this is quite significant, especially when running large models with long context windows. That VRAM could otherwise be used for additional context, larger KV cache, draft models, or other features.

CPU-based image processing is an alternative, but in my experience it creates substantial overhead and is much slower.

For this reason, I currently often disable mmproj completely for Qwen and Gemma models: keeping it on the GPU costs too much VRAM, while processing images on the CPU is too expensive.

A low-memory GPU implementation would provide a useful middle ground:

  • fast GPU-based image processing

  • very low additional VRAM usage

  • more VRAM available for context and KV cache

  • no need to choose between fast vision processing and large context sizes

For users running modern multimodal models on GPUs with limited VRAM, this could make a substantial difference.

Possible Implementation

HyperQwen may provide a useful reference implementation. Its vLLM-based implementation seems to use a sliding-window approach for image processing, allowing the processing to remain GPU-accelerated while using only a very small amount of VRAM.

The related llama.cpp discussion also appears to cover approaches in this direction:

ggml-org#20246

I assume that a fully model-independent implementation may not be straightforward, since projector architectures differ between model families.

Still, even an implementation initially targeting a few common architectures such as Qwen and Gemma could already be very useful, and might provide a basis for a more generic solution later.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions