Skip to content

local Ollama depend on 'llama cpp' #75

Description

@blak-Dev

Please consider add option for 'llama cpp' as offline provider
https://github.com/ggml-org/llama.cpp

Launch OpenAI-compatible API server

llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
why extra mile for Ollama

Activity

  1. DevMando commented on Aug 19, 2026

    @DevMando
    Owner

    @blak-Dev Thanks for filing this. You're the first person outside the project to open an issue here, so it honestly made my week.

    Fair question, and there wasn't a grand reason. When I started MandoCode I was pretty new to the AI space, saw that Semantic Kernel had an Ollama connector, and just ran with it. It worked, so it stuck. I've spent a lot more time down in the lower layers since then, and this has been sitting in the back of my head for a while, so your timing is good.

    Here's where it actually stands. BuildKernel calls AddOllamaChatCompletion, which talks to Ollama's native /api/chat instead of /v1/chat/completions, and model validation POSTs to /api/show. So pointing ollamaEndpoint at llama serve won't work today. The execution settings and token counting are Ollama-typed too. It's a provider abstraction rather than a config flag. Good news is Microsoft.Extensions.AI is already referenced, so the IChatClient seam is half paid for, and one OpenAI-compatible path would cover LM Studio, vLLM and hosted gateways at the same time. I'm also considering and looking into updating the project entirely to Agent Framework, but that's another conversation.

    One thing worth mentioning that's in your favor: Ollama maps HF GGUFs onto a Go template it guesses from the GGUF metadata, while --jinja uses the model's real Jinja template. MandoCode leans really hard on tool calls, and a remapped template can render chat fine while quietly mangling tool call syntax. So llama.cpp might actually be the better backend here, not just another option.

    Anyway, I'm picking this up as the next change to the provider layer, thanks for the feedback! cheers

  2. self-assigned this
    on Aug 19, 2026
  3. DevMando commented on Aug 23, 2026

    @DevMando
    Owner

    Quick update, my comment above is already out of date.

    That Agent Framework migration I said was another conversation? Done and merged.
    Semantic Kernel is completely out of the codebase now.

    So the BuildKernel / AddOllamaChatCompletion stuff I described doesn't apply
    anymore. There's exactly one place left that constructs an Ollama client, and it
    hands it straight to the agent as an IChatClient. Dropping in an OpenAI
    compatible client is a construction point now, not an orchestration rewrite. Nice
    side effect of having done that migration for other reasons.

    What's left is the stuff around it. Model listing and validation still hit
    /api/tags and /api/show. The context window is the interesting one. Right now
    MandoCode stamps num_ctx onto every /api/chat request, which is what makes the
    setting work no matter who started the daemon, and llama.cpp's server fixes
    context at launch instead. So under llama.cpp contextLength stops being something
    MandoCode sets per request and becomes a flag you pass when you start the server.
    Not a blocker, but it changes what the setting means, and I'd rather say that out
    loud than let it quietly do nothing.

    So still a provider abstraction rather than a config flag, just a much smaller
    one than when I wrote that.

    The template thing I brought up is still what I keep chewing on. Tool calling is
    where this actually gets decided, and I'd rather test that properly than assume
    it works.

    Anyway, still next up in the provider layer. I'll post here when there's
    something you can actually run. cheers

  4. added
    consideringDesign question I'm still weighing. Open on purpose, not stalled.
    providerModel provider / backend support
    on Aug 27, 2026
  5. DevMando commented on Aug 27, 2026

    @DevMando
    Owner

    Update on this one. I'm pausing it for consideration rather than picking it
    straight up, and I've labelled it so it's clear that's deliberate.

    The reason is that it isn't really a llama.cpp feature, it's a provider layer.
    Once MandoCode can talk to something other than Ollama, the same seam covers LM
    Studio, vLLM, Azure Foundry and whatever else turns up later. If I build it I'd
    rather build it as that, properly, than bolt on one backend and end up doing the
    work twice.

    Which makes it a bigger piece than the original ask, and a long term commitment
    to maintaining more than one provider. MandoCode is still very much being built
    out, so I want to be sure before I take that on.

    Not dropping it and not a no. I just want to plan it as an enhancement in its own
    right rather than squeeze it in.

    One question while I'm thinking it over. Would local only cover what you're
    after, or were you hoping this opened the door to hosted OpenAI compatible
    endpoints too? cheers

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

consideringDesign question I'm still weighing. Open on purpose, not stalled.enhancementNew feature or requestproviderModel provider / backend support

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions