Skip to content

About

A small self-learning coding agent for your own model — any OpenAI-compatible server, local (LLMTray, LM Studio, Ollama, vLLM) or cloud. Single static Go binary, plan/auto modes, skills, permissions, Claude-Code-style TUI.

Topics

Resources

Contributing

Stars

4 stars

Watchers

0 watching

Forks

Repository files navigation

ipsupport-code

ci release license go

ipsupport-code — analyze, fix, test, report

Website: ipsupport-llc.github.io/ipsupport-code

A small, lightweight self-learning coding agent for your own model — any OpenAI-compatible server, local or cloud — in a single static binary, with a built-in risk model that stops a dangerous tool call for you. It drives that model through a reason → act → observe loop over a handful of fat tools (file, run, git, web, calc), recovers from tool errors using lessons it learned on past runs, and — after each task — reflects and writes new lessons to disk so it actually gets better over time.

Local by default, but the same loop runs any OpenAI-compatible provider (OpenAI, Anthropic, Grok, Groq, OpenRouter, Z.ai) — switch in one command — and it can delegate a task to a sub-agent on a different model, so a local agent can hand a review to a frontier one and compare. See Providers and Sub-agents.

The interactive UI is a Claude-Code-style TUI (Bubble Tea): live-streamed tool calls and observations, markdown answers, syntax-highlighted diffs, a plan / auto mode toggle, non-stealing approval prompts, and a status bar showing the model, context usage, and tokens. Piped/non-interactive runs fall back to a plain line REPL.

It will not match a frontier model. It's built for micro-tasks on your own machine, with your own model, under a permission policy you control.

Not affiliated with Anthropic's Claude Code; the name reflects a similar terminal-agent UX.

At a glance

  • 🧰 Fat tools — file · run · git · web · calc, behind a permission policy + workspace jail
  • 🧠 Self-learning — distils lessons from each run and recovers from repeat mistakes
  • 🎯 Goals — /goal pursues a multi-turn objective; a judge re-feeds it until it's actually met
  • 🌐 Any model — local first (LLMTray on a Mac, LM Studio, Ollama, vLLM — keyless local servers included), or any OpenAI-compatible provider
  • 🤝 Sub-agents — delegate/fan-out across other models or local CLI agents (codex/claude/…) and merge the results
  • 🛎️ Steer live — /steer <note> nudges a running task mid-flight (and /btw <question> asks a quick side question), without stopping it (esc still cancels)
  • 🛡️ Built-in risk guardrail — a local model inside the binary scores every tool call; a flagged one stops for you even when the policy allows it, and it learns from your y/n (Risk check)
  • 🪶 Lightweight — a ~2.7 KB system prompt and a ~5 KB tool catalog, about 2,000 tokens before your task; one ~20 MB binary, no runtime
  • 🔒 Read-only runs — -read-only answers a review or a question with nothing allowed to change
  • 💰 Guardrails — /budget spend cap per run · /diff to review what the agent changed
  • 🧩 Plan/auto modes · ⏪ /rewind · 🔌 MCP · 📦 skills · ♻️ self-updating

New here? Install below, then jump to Quick start.

Install

One-liner (macOS/Linux) — auto-detects your platform, verifies the SHA-256, installs to ~/.local/bin, and prints how to run it (and how to add it to PATH if needed):

curl -fsSL https://raw.githubusercontent.com/ipsupport-llc/ipsupport-code/main/scripts/install.sh | sh

That installs the latest nightly; append -s -- latest for the newest stable release, -s -- v0.22.0 for a specific tag, or a second arg for a custom path.

Windows (PowerShell) — installs to %LOCALAPPDATA%\Programs\ipsupport-code, verifies the SHA-256, and adds it to your user PATH:

iex (irm https://ipsupport-llc.github.io/ipsupport-code/install.ps1)

That's the nightly; for a channel/tag use the scriptblock form (so it can take an arg): & ([scriptblock]::Create((irm https://ipsupport-llc.github.io/ipsupport-code/install.ps1))) latest.

Both installers also add the short name ipco beside it (a symlink; on Windows a one-line ipco.cmd), unless something else already owns that name.

Homebrew (macOS/Linux) — the latest stable release, with ipco beside it:

brew install ipsupport-llc/tap/ipco

Homebrew then owns the install: update points you at brew upgrade ipsupport-code instead of replacing the binary.

Or download the archive for your platform from the latest release (or the rolling nightly): darwin-arm64 (Apple Silicon), darwin-amd64 (Intel Mac), linux-amd64, linux-arm64, windows-amd64, windows-arm64. Each release also ships checksums.txt (SHA-256).

tar -xzf ipsupport-code_*_darwin-arm64.tar.gz   # .zip on Windows
chmod +x ipsupport-code
# macOS may quarantine a downloaded binary:
xattr -d com.apple.quarantine ./ipsupport-code
./ipsupport-code -version

Or build from source (pure Go, CGO_ENABLED=0, cross-compiles from any host):

make build     # host binary at dist/ipsupport-code
make release   # every target into dist/
go install github.com/ipsupport-llc/ipsupport-code/cmd/agent@latest  # installs as `agent`

Providers — local or external

The agent and tools are model-agnostic, so you can point them at any OpenAI-compatible API and switch in one command — run a strong cloud model for a hard task, then drop back to your local one:

/ai                       pick a provider from a list (built-ins: openai, anthropic, grok, groq, openrouter, zai)
/ai key openai sk-…       add an API key for a built-in provider (one command)
/ai add mylab https://api.lab.co/v1 llama-3.1 key=sk-…   add a CUSTOM provider + key in one step
/ai openai                switch to it   ·   /ai local   back to your local server
/model                    pick from the server's models (↑↓, type to filter)   ·   /model gpt-4o   switch directly

Pasting a key on Linux? A middle-click paste is swallowed by the mouse-wheel scroll tracking — use Ctrl+Shift+V (bracketed paste), which always works.

Add a provider by hand. Any OpenAI-compatible endpoint works as a first-class provider straight from the config — no command needed, and no API key required for a keyless local server (Ollama, vLLM, a second LM Studio). Add it under providers and point provider at it:

{ "provider": "ollama",
  "providers": { "ollama": { "base_url": "http://localhost:11434/v1", "model": "qwen2.5-coder:7b" } } }

It's then active at launch and switchable via /ai ollama and /config.

Switching keeps your session, tokens, and mode. Keys live in ~/.config/ipsupport-code/config.json (written chmod 600) or fall back to the env var (OPENAI_API_KEY, ANTHROPIC_API_KEY, XAI_API_KEY, GROQ_API_KEY, OPENROUTER_API_KEY, ZAI_API_KEY). /config opens an interactive settings panel: ↑↓ to move, Enter to change the selected row, esc to close — changes apply and save as you make them, no hand-editing JSON.

  • Connection edits the provider in use, the local server included: the provider and model are picked from a list (↑↓ to scroll, type to filter; a model the server doesn't list can be typed), its address (host, port, path), API key (typed masked; empty keeps it, ctrl+d removes it), server type (LM Studio or plain OpenAI-compatible) and context window (any size; 0 = auto-detect). A built-in provider's address can point at a proxy or your own gateway; empty puts it back.
  • Providers adds or edits an OpenAI-compatible endpoint, and removes one picked from the saved ones (confirmed with a second Enter; the one in use falls back to local).
  • Tuning values — temperature, top_p, output cap, idle timeout, retries, and the goal judge's output cap — are typed; the rest cycle in place. Values are edited like a shell line: ←/→, home/end, ctrl+u.
  • While a task runs, reasoning (also /reasoning <level>), temperature, top_p, the output caps, retries and loop detection take effect from the model's next reply and are saved when the task ends; other changes are staged until then.
  • A setting this project's .agent/config.json also sets is marked: that file wins at the next start.

Sub-agents

Define a profile — a named model — and the agent gains an agent tool: it can delegate a self-contained task to a sub-agent on that model and use the answer. A second opinion, another model's strength, or the same code reviewed across 2–3 models at once. Run your main loop on a local model, then have it fan a review out to a couple of frontier models and merge the findings.

A profile is the only way to delegate (and so the curated list of what the assistant may spawn). Build one interactively — provider → model list → name — from /config, or on the command line:

/agents add grok   openrouter x-ai/grok-4.3     # omit the model to list them
/agents add architect openai gpt-4o
/agents                                          # list profiles
/agents rm grok                                  # remove one

Then just ask — e.g. "review internal/tool across grok and architect, then merge." The assistant calls agent(profile, task, dir): dir (optional, ~ expanded) points a sub-agent at another project — it gets its own jail there, so it can work on ~/other-repo without leaving the rest of your disk exposed. A pure fan-out of several profiles runs in parallel, each on its own live status line; the main model then merges the results into one answer.

Safety: every spawn asks for approval (even local ones cost compute) until you relax it with /permissions agents on. Sub-agents read/write files and use git but only run shell commands if you enable it with /agents exec on. They inherit the current plan/auto mode, can't spawn their own sub-agents (depth 1), and their tokens are recorded in /usage like any other.

External CLI agents

Locally installed CLI coding agents (Codex, Claude Code, Muse, aider…) can be sub-agents too — registered as external profiles:

/agents add-tool                # scan PATH: which known CLI agents are installed
/agents add-tool codex          # one word — codex/claude/gemini/qwen/aider/goose/opencode/muse
/agents add-tool mytool mytool --headless {task}   # any other tool ({task} = where the task goes)

The assistant delegates through the same agent tool; the CLI runs in the target dir, and the assistant gets back the tail of its output plus a git diff --stat summary (review the full patch with /diff). Use the CLI's non-interactive mode (exec / -p / --message) — an interactive launch just hangs until the timeout (15 min default, timeout per profile in config.json). Muse (Meta's Muse Code) runs as muse exec {task} and keeps its own approval and sandbox on in that mode; to launch it with other flags, use the full form (/agents add-tool muse muse exec <flags> {task}).

⚠ External agents run outside the sandbox: their own tools, their own permissions, no policy jail, and /rewind can't see their edits. That's why every launch asks its own approval (external CLI agents category) — even when ordinary spawns are set to allow; a on the prompt relaxes it for the session.

Background jobs (fire-and-forget)

Any sub-agent — LLM or external CLI — can run as a detached background job: the assistant adds background=true, the call returns at once, and it keeps working while the job runs. When the job finishes you see a ✓ job #N notice, and the result is folded into the assistant's next step — even mid-task, via the same between-steps seam as /steer, so a job that lands while the assistant is working doesn't wait for the next task to be noticed. /jobs lists them (each running one shows ⚙ running <elapsed> plus ↳ <its latest output line> (<age> ago) for an external agent — a long gap hints it's stuck, not working), /jobs result <id> prints one in full, /jobs kill <id> cancels. Jobs survive their parent task (and esc) — perfect for a long codex review running while you continue in the main loop.

Steering a running task (/steer) and asking on the side (/btw)

esc cancels a task. Two lighter touches don't stop it:

  • /steer <note> folds a note into the running task — it lands on the very next step to redirect the work ("/steer the loader is in internal/config, not cmd"), without waiting for the task to finish or queuing as a follow-up. The note is pinned above the input for the rest of the run. Typed while idle, it steers the next task you start.
  • /btw <question> asks a quick side question while the task keeps going — it's answered in one turn, no tools, from the live conversation, then work resumes ("/btw what test framework is this repo using?"). The answer doesn't steer the task; it's just for you.

Prompt templates (/snip)

Save a prompt you reuse and pull it back when you need it. /snip save <name> stores your last prompt under a name (or /snip save <name> <text> to store text directly); /snip <name> drops that template into the input box so you can tweak it and hit Enter — not auto-sent. /snip list shows them, /snip rm <name> deletes. Templates are stored globally (~/.config/ipsupport-code/snippets.json), so they follow you across projects, and Tab completes both the subcommands and your snippet names.

Updating

The binary updates itself in place from GitHub Releases:

ipsupport-code update            # update on the current channel
ipsupport-code update nightly    # switch to the nightly channel (saved) and update
ipsupport-code update stable     # switch back to stable

/update does the same from inside the REPL. On startup it does a quick, best-effort check and prints a one-line notice if a newer build is out on your channel (stable by default — set "channel": "nightly" in ~/.config/ipsupport-code/config.json, or switch with update nightly). Local dev builds are never nagged. To silence the startup check on a pinned or package-managed install without going fully offline, set ipsupport-code config set update_check false.

Working directory. Launched from a parent dir (or ~)? /cd <subdir> points the session at your project — relative file/run/git paths resolve there, and sub-agents inherit it as their default dir, so you set the path once instead of repeating it. It stays inside the workspace jail.

Offline? /offline on cuts the agent's OWN internet use — the web tool refuses with a clear "no internet" message and the startup update check is skipped. Your configured model connection is untouched either way — if it's a local server on localhost it needs no network anyway; if you pointed it at a remote endpoint, that traffic still goes out. /offline off re-enables it.

Usage statistics and rating

Anonymous usage statistics help decide what to work on. They are on by default for a new install — first-run setup says so — and an update never turns them on for an install that existed before: it is written off.

  • What is sent: one report per day, at the first launch after that day ends (nothing runs on a timer). It holds the version, OS and its version, CPU architecture, Apple chip (Apple silicon only), memory size, language, how many times each feature was used (tasks, tool_calls, goals, subagents, external_agents, mcp, skills, plan_mode, local_model, cloud_model), the coarse family of models used (qwen, nemotron, claude…) and a random install ID.
  • Never sent: code, prompts, commands, tool arguments or output, file names or paths, model names, API keys, provider URLs.
  • /telemetry shows how the last send went (accepted, refused and why, or failed), the days waiting and exactly the reports they will send. /telemetry off (or ipsupport-code config set telemetry false) turns it off and deletes the install ID and every unsent counter; /telemetry reset makes a new ID.
  • Nothing is recorded under DO_NOT_TRACK=1 or in /offline mode, from a piped first run that never saw setup, or by a development build. A workspace's .agent/config.json can't turn it on.
  • The first report goes only after the first day ends, so it can be turned off before anything leaves the machine. The server's side — what it keeps and for how long — is ipsupport-api's telemetry spec.

Rate it from inside the agent: /rate 5 works great with my local model (add --name <you> to sign it). Ratings are moderated, then shown on the website. After the first day a single dim line may suggest it — at most once every two weeks; /rate never stops that.

Quick start

  1. Start a local server with a tool-calling model loaded (e.g. qwen2.5-7b-instruct). On a Mac we recommend LLMTray — a menu-bar app running mlx-lm natively on Apple Silicon; LM Studio, Ollama and vLLM do the same job anywhere. No local model? Skip this — the next step asks.

  2. First interactive run asks whether you have a local model server running. If yes, it looks for a running LLMTray (:8765) or LM Studio (:1234) and offers the one it finds as the server URL, then asks for the API key (blank for a local server) and model, and confirms the connection. If no, it asks for a provider name — a built-in (openai, anthropic, grok, groq, openrouter, zai) uses that vendor's real endpoint with just a key; any other name is treated as your own OpenAI-compatible endpoint (a self-hosted gateway, LiteLLM, a proxy…) and additionally asks for its Base URL — so a custom endpoint is never silently pinned to the wrong vendor's URL. Either way it saves to ~/.config/ipsupport-code/config.json (re-run with -init).

  3. Run a one-shot task, or open the REPL:

    ./ipsupport-code "use calc to compute (1234*9)+sqrt(2)"
    ./ipsupport-code                      # interactive TUI
    ./ipsupport-code -C ~/proj "summarize what main.go does"

Which session? Memory is saved per workspace (the -C dir, default the current one) and per agent name (default ipsupport-code). A one-shot or piped run silently continues the saved thread; the interactive TUI opens a navigable chooser (↑↓ to pick a saved session, enter to open, d to delete, esc for the newest) when any exist. To start clean or pick a thread from the CLI:

Flags go before the task text — anything after it is part of the task, not a flag, and the agent says so rather than quietly sending it to the model:

$ ipsupport-code "do the thing" -override provider=openai
error: -override provider=openai must come BEFORE the task text — written there it was read as part of the task, not as a flag

To point a single run at a different provider or model without touching config.json (the provider must already be configured):

./ipsupport-code -override provider=openai "task"                    # this run only
./ipsupport-code -override provider=openai \
                 -override providers.openai.model=gpt-5 "task"       # …and its model
./ipsupport-code -new "fresh task"              # ignore the saved session
./ipsupport-code -session review "look at X"    # start a named thread
./ipsupport-code -session review -resume        # continue it
./ipsupport-code -session review -new           # start it over (overwrites)

A named session is an explicit choice, so relaunching one without saying which you meant is an error rather than a silent guess:

$ ./ipsupport-code -session review
error: session "review" already exists — -resume to continue it, or -new to start over (overwrites it)

Read-only runs — for a review, or a question about the code, with nothing allowed to change:

./ipsupport-code -read-only "review the diff against main"

File writes, git changes and MCP servers are refused. Commands run without asking, inside the OS sandbox (Landlock on Linux 6.7 or later, Seatbelt on macOS, whatever sandbox is set to) with the workspace read-only like the rest of the disk: only temp dirs and the user cache dir are writable, and there is no network. The git tool's own git runs there too, since a status or diff runs the filters the repository's config names. Where there is no such sandbox, or the workspace or its repository sits in a temp or cache dir, commands and the git tool are refused too, and the run says why. The session is neither restored nor saved, and nothing is learned from it — not even from an approval answer; a run whose own files (trace, usage, skills, lessons, settings) would land in the workspace or its repository doesn't start. A repository nested inside the workspace with its git dir kept elsewhere is not looked for. On Linux a command can still change a file's mode and times (Landlock doesn't cover them), never its contents.

Each named session keeps its own saved thread, its own standing goal and its own input history, so two of them can work the same checkout at once — a local model in one terminal, a cloud model in another — without stealing each other's goal. Learned facts and lessons stay shared: those describe the project, and both sessions benefit from what either one learns.

Logs are per workspace and per session:

~/.config/ipsupport-code/agent-test.log          # -C test
~/.config/ipsupport-code/agent-test-remote.log   # -C test-remote
~/.config/ipsupport-code/agent-test-cloud.log    # -C test -session cloud

Named by the directory's basename, so you can tail one without looking up a hash. (Two checkouts with the same basename share a log — unlike the state directory, where a collision would mix two projects' memory rather than two logs.)

Every shared state file is written under a cross-process lock, and the ones both runs add to — learned facts, lessons, the usage ledger, input history — are merged under it rather than overwritten, so a second run never erases what the first learned.

In the TUI, /sessions lists/switches/deletes threads, /new <name> starts a fresh named thread (the old one stays in /sessions), and /new clears the active one.

Modes: auto vs plan

Toggle with shift+tab (or /plan / /auto); the current mode shows at the bottom of the screen.

  • auto (default) — the agent executes the task with tools.
  • plan — read-only: it investigates and proposes a numbered plan, and every state-changing tool call is blocked at the engine, so it can't touch anything until you accept.

When a plan-mode task finishes with a plan, you get a one-key handshake: enter to accept — switches to auto and immediately executes the plan — or esc to keep planning (stay in plan mode; type to refine). Toggling the mode mid-task applies at the next turn, never mid-run (the running agent reads the mode live).

Permissions

Mutating actions ask for approval by default. The prompt goes modal: y approves, n denies, a allows the whole session (see below), ↑/↓ toggle the explicit Yes/No, and every other key is ignored outright — nothing leaks half-typed into the chat while it's waiting. esc drops back to typing (to steer or queue something else first); the approval stays pending either way, and answering it always lands you back where you were. A non-overridable deny floor (rm -rf, sudo, secrets, .git, .env, …) is always enforced.

Grant it for the session. Press a on any prompt to stop asking about that whole kind of action (file changes, shell, git, sub-agent spawns, MCP) for the rest of the session — in memory only, never written to config, cleared by /new and /clear. It's the "yes, and don't ask again for now" you reach for mid-task without loosening anything permanently.

For a permanent relaxation instead, /permissions files on auto-allows non-destructive file ops in the workspace (the deny floor still applies); same for /permissions run on. That choice is saved to the workspace config.

OS sandbox (opt-in). The permission engine decides whether a command runs; an OS-level sandbox can additionally confine what it can touch once it does — so even an allowed command can't write outside the workspace. Turn it on with config set sandbox auto (picks the platform's mechanism: Seatbelt on macOS, Landlock on Linux — kernel 5.13+), or name one explicitly (seatbelt / landlock). Shell (run) commands then execute inside the kernel sandbox: writes are confined to the workspace (plus temp dirs and the user cache, so build tools keep working), the network follows /offline (Landlock needs kernel 6.7+ for that part), reads and everything else are unchanged. It's off by default, and on a platform/kernel without a supported sandbox commands run unconfined as before. External CLI agents are not sandboxed (they run outside it, as documented above).

Risk check

A small local classifier scores every tool call for how dangerous it looks. A call it flags — or one it could not read in full — stops for you even where the permission policy would have run it without asking: an allow glob, a run.default of allow, an earlier a on a prompt. Your answer is the override: y runs it, n refuses it (the model is told), and a stops asking about that kind of call for the rest of the session, flagged ones included. An a given at an ordinary prompt covers only the unflagged ones. An external agent's or an MCP server's launch command comes from config, not from the call the model made — a checkout's own file can set it — so it is scored too, as it starts.

Turn the asking off with the risk check row in /config (saved as risk_check), or for one run with -skip-risk-check. -skip-permissions turns it off too: nothing is asked at all. A project's own .agent/config.json can turn it on but not off — it is what stops a flagged command in a checkout whose file allows everything. On or off, the scorer logs what it thought next to what the policy did:

msg="risk shadow" tool=run action=shell risk=1.00 top=credential_access
  labels="credential_access=1.00 sandbox_escape=1.00" policy=allow
  disagreement=allowed-but-flagged call="run shell command=cat ~/.ssh/id_rsa"
msg="risk shadow: scored 1 call(s), 1 over 0.50, 1 disagreed with the policy"

You see it where it matters, not only in the log. A call the scorer is unhappy about carries its suspicion on its own line, and again at the approval prompt — which is the moment it is worth anything, since your answer there is the label it learns from:

  ⚙ run shell cat ~/.ssh/id_rsa   ⚠ possible credential access (0.98)
  ⚠ approve run: cat ~/.ssh/id_rsa   ⚠ possible credential access (0.98)
    y approve · n deny · a allow all shell commands this session, flagged ones too

Nothing is shown for a call it had nothing to say about — a number on every routine line is how a risk signal gets tuned out.

The disagreement column is the whole product. allowed-but-flagged is what the check stops; gated-but-unremarkable is what the policy asks about that the scorer would not have. Tail them with IPS_LOG=debug, and turn scoring off entirely with IPS_RISK=off.

It is a hashed-feature linear model — sigmoid(Wx+b) over word and character n-grams, six labels (destructive, sandbox_escape, credential_access, network, external_side_effect, safe). Inference is pure Go with no dependency of any kind: 768KB of float32 embedded in the binary, ~28µs per call. Training is offline and separate (scripts/train_risk.py, stdlib only); its only output is the weights file.

A shell line is scored by its parts. Read as one text, a long harmless tail diluted a destructive head: rm -rf quotesdemo && cat > REPORT.md <<'EOF' … came out at 0.07. The line is now cut by the rules of the shell it runs in — sh off Windows, PowerShell (or cmd when run.shell says so) on it (internal/shellsplit) — and its code (commands and operators, without heredoc bodies, here-strings or comments, which are data) each group of commands joined by | && || (a pipeline's danger can be in the joining: curl … | sh) and each command in it are scored; the call's risk is the highest. A heredoc that is expanded or fed to a shell or an interpreter (sh <<EOF, python3 - <<EOF) is code, not data. That line now scores 0.89, and a note written through a heredoc no longer fires on what the note says. Paths into the workspace are read as the project's own (/app/data.csv in a workspace at /app scores as ./data.csv): the scorer knows where the workspace is, the model can't.

Swap the model without rebuilding: IPS_RISK_MODEL=/path/to/model.bin. The file carries its own feature config and label names, so a model with a different feature space or a different set of labels loads unchanged.

On held-out paths — a fifth of every path class, never seen under any verb:

label precision recall
destructive 0.81 0.97
sandbox_escape 0.96 0.84
credential_access 0.93 0.95
network 0.99 0.99
external_side_effect 0.99 0.99
safe 0.96 0.88

On the headline score — what the log shows and a gate would use — that is 7.5% false alarms on ordinary calls and 5.3% missed risky ones (chains of commands share their members' held-out paths, so an unseen path's mistake repeats in them). Most false alarms are deleting build output under a name never seen (artifacts, local.properties); most misses sit just under the line on a held-out secret or system path. On the honest set of real agents' commands (scripts/risk_eval.jsonl), the number that decides whether a model ships: 169 of 198 risky commands flagged, 42 of 501 ordinary ones — against 124 and 70 before lines were scored by their parts. The trainer prints both, and names every false alarm.

It knows what makes another agent's or an MCP server's launch dangerous (see Risk check): the flags that turn an agent's own approvals and sandbox off — codex --yolo, claude --dangerously-skip-permissions, qwen --yolo, muse exec --disable-sandbox, copilot --allow-all-tools, opencode run --dangerously-skip-permissions, cursor-agent --force, this agent's own -skip-permissions — and a library slipped into a program through LD_PRELOAD or DYLD_INSERT_LIBRARIES. The same agents with their ordinary flags, and servers started with a key in their environment — any of gitleaks' vendors, named but never valued — stay quiet, and so does an agent run with its approvals off but its sandbox on (codex exec --full-auto, muse exec --approval-mode never, gemini --yolo, whose sandbox yolo turns on by default): that is contained. Gaps: an agent it has never seen (its safety-off flag is just an unknown word), and a variable that names code to load (PYTHONPATH, NODE_OPTIONS), which a launch line shows without its value and agents point outside the workspace all day.

The feature vector is L2-normalized, which is what makes a score mean the same thing for a long call as a short one. Without it a repeated feature accumulates: the same command with a 380-character path had Sum(v²) = 426490 against 104, and a logit of 1523 against 14.5 — putting it beyond the reach of the cap on local corrections, so it could never be corrected at all.

The vocabulary is vendored from upstream, not invented: build-output names come from github/gitignore's 309 templates (CC0), and the words that mark a secret — token, api, key, secret, plus ~130 vendor names — from gitleaks' 222 rules (MIT). scripts/fetch_risk_vocab.py is the only script that touches the network; its output is committed, so generating the dataset, training, and make build all stay offline and deterministic.

Upstream needs judgement applied, and the judgement is recorded in scripts/risk_vocab.json rather than hidden: "do not commit this" is not "safe to delete", so Makefile, README.txt and app/config/parameters.yml are vetoed out of the safe half, along with names that are generated in one ecosystem and hand-written in another (docs, public, lib). The two failure modes are not symmetric — missing a build directory costs one score that reads high; calling docs build output teaches the model that rm -rf docs is routine.

The dataset is built compositionally — every verb crossed with every class of argument — so the verb carries almost no information and the argument carries all of it. cat README.md scores 0.00 and cat ~/.ssh/id_ecdsa scores 1.00; so do less, wc -l, xxd and od -c on the same two files. That took three attempts to get right (see the comment at the top of scripts/gen_risk_dataset.py), and it is the whole difference between a risk signal and a list of scary words.

network is informational: reaching the internet is a property, not a danger, so it is reported but excluded from the headline number — otherwise fetching a documentation page scores 1.00. Which labels are informational is declared in the model file, so a replacement model decides for its own.

It learns from your approvals

The base weights are trained on synthetic data and never change. What a run learns goes beside them, per workspace — and it learns from the only ground truth the agent gets for free: you answering an approval prompt. Two answers carry information, and they are exactly the two the shadow log already counts as disagreements:

it flagged the call and you approved a warning you didn't need — calls like it stop being flagged here
it stayed quiet and you refused a miss — calls like it start being flagged here

What your answers change is whether a call gets a warning in this workspace, not what the model says the call does. An approved git push still reads external_side_effect 0.95 in the log and in /risk — it does push somewhere — but after a few approvals it stops interrupting you here. The log shows both numbers: risk after this workspace's corrections, base before.

An approval of something it also thought was fine teaches nothing; that would be learning from its own output. Neither does anything that isn't an answer — esc, a killed background job, a closed stdin all deny the call, and none of them is a person judging it.

Every agent scores its own calls against its own policy, sub-agents included. A delegate without its own scorer inherits whatever assessment is on the context it was handed — the parent's agent.spawn — and then a refusal of the delegate's own file write gets recorded against the spawn.

A refusal can mean "not now" or "I'll do it myself" as easily as "that is dangerous", so one answer never flips the model. A borderline score settles on the first correction; a confident wrong one takes five or six consistent answers. Corrections that stop being repeated decay away, and /risk reset drops them all.

/risk shows what happened and what was learned. Every answer is also appended to risk-feedback.jsonl in the workspace's state directory, in the shape scripts/risk_dataset.jsonl uses — as a verdict ("verdict": "approved", with the labels the model scored), not as labels. An approval says the call was acceptable, not what it does, so the trainer skips these rows. To teach the base model something, label a row yourself — write its labels and set "source": "manual" — and fine-tune from the existing weights:

python3 scripts/train_risk.py --from internal/risk/model.bin   ~/.config/ipsupport-code/state/<workspace>/risk-feedback.jsonl

Rows written before v0.61, which recorded an approval as ["safe"], are skipped the same way.

Two runs of the trainer on the committed dataset produce a byte-identical model.bin, so the shipped model is reproducible from the files in this repo.

What it is not. It does not replace the permission policy or the sandbox. The policy is still better at the extremes — it denies rm -rf / outright, whatever the allow-list says. And learning needs the prompt: with -skip-permissions nobody is asked, so nothing is answered and nothing is learned.

Skills

On-demand instruction packs — the user-extensible version of guides-on-demand. Only an enabled skill adds a single line to the system prompt; the model loads a skill's full instructions on demand, so the base prompt stays lean no matter how many you install. Eight curated skills ship in the binary (test-first, debug-systematically, git-flow, research-first, minimal-code, review — multi-model review via sub-agents — subagents, how to delegate and fan work out, and plan, a .agent/plan.md checklist so multi-step work resumes itself), seeded disabled so you opt in. Built-in skills refresh on upgrade unless you've edited them.

/skills                       list installed skills (on/off)
/skills on git-flow           enable one
/skills install <url|git>     add a .md by URL, or every skill in a git repo
/skills remove <name>

MCP servers

Connect Model Context Protocol servers and the agent gains their tools — but through one proxy tool, not by dumping every server's schema into the prompt (which would swamp a small model's context). Add them to ~/.config/ipsupport-code/config.json:

"mcp_servers": {
  "fs":     { "command": "npx", "args": ["-y", "@modelcontextprotocol/server-filesystem", "/some/dir"] },
  "gh":     { "command": "github-mcp-server", "env": { "GITHUB_TOKEN": "…" } },
  "remote": { "url": "https://mcp.example.com/mcp", "headers": { "Authorization": "Bearer …" } }
}

Two transports: stdio (a local subprocess via command) and HTTP (a remote url, with headers for auth — e.g. Authorization: Bearer). /mcp lists the servers and their tools. The model uses the mcp tool — list to discover, schema to see a tool's inputs, call to run one. Servers connect lazily on first use; every mcp call asks for approval (it's external code). Sub-agents don't get MCP.

When a local model misbehaves

Small/“thinking” local models can loop in their own reasoning or over-think. The levers, in order:

  • /reasoning low (or minimal/off) — trims the model's reasoning. There's no universal API for this, so it's per-model and stored in its provider's own shape (reasoning_effort for OpenAI, reasoning:{effort} for OpenRouter, chat_template_kwargs:{enable_thinking} for Qwen/LM Studio). /reasoning low on the current model writes the right shape for known providers; for others set it raw in config.json under reasoning (keyed by <provider> or <provider>/<model>).
  • A runaway turn is auto-stopped. A single turn generating far past the context window (looping) aborts with a clear message instead of streaming for minutes. Press esc to cancel anything sooner.
  • /reflect off — skip the post-task learning pass if it's where a weak model loops (the status shows “task done — distilling lessons” so you can tell the task itself already finished). Or /reflect <profile> to run learning on a stronger model, and /reasoning reflect low to give the learning pass its own (leaner) reasoning setting.
  • /skills on plan — for long multi-step work, the model keeps a checklist so it resumes instead of drifting.
  • Fine-tune it on the tool schema. If a specific local model keeps garbling the {"action":...,"params":...} shape (the registry recovers several common ones automatically, but not all), see FINETUNING.md for a from-scratch guide, including where to mine real training examples you already have on disk.

Rewind

/rewind opens a list of the session's steps — pick one (↑↓, enter) to roll back to before it ran: every file that step (or its sub-agents) changed is restored to its prior content, files it created are removed, and the conversation is trimmed back. Checkpoints are taken at the start of each turn, so you can rewind no matter how a turn ended (finished, esc, a loop, an error). Shell commands, git, and network calls can't be undone — only files and the chat. (REPL: /rewind lists, /rewind <n> applies.) Snapshots live for the session.

Goals

A goal is a multi-turn objective — not a single turn. /goal <text> sets one and starts pursuing it: the agent works, and when it thinks it's finished a judge (a separate model call) decides whether the goal is actually met. If it isn't, the goal is re-fed to the agent — kept in focus, with the gap the judge named — and it keeps going. This repeats up to a TTL (/goal ttl <n>, default 6 re-feeds) before it gives up, on top of the usual esc / stuck / runaway guards and a hard step cap.

The goal is a first-class, persisted object: it lives in the workspace's state directory (goal.json), so it survives a restart — an unfinished one offers to resume once on the next start (not every time), and a completed goal clears itself. /goal shows the standing goal and its status; /goal go resumes it; /goal clear drops it; /goal off disables the loop (a goal then runs as a single pass). Plain tasks (anything you type that isn't a goal) run as one pass with no judge overhead — only an explicit goal gets the loop.

The judge defaults to done on any unparseable reply, so a confused model can't trap the agent in the loop. The whole thing is gated on real progress: a turn that calls no tools is never judged or re-fed. If a re-fed model then finishes without doing any work, it gets one push to act before the loop gives up (turn it off with config set goal_nudge false). And when the loop stops without the judge ever confirming success, it says so plainly — "goal not confirmed complete, /goal go to keep pushing" — instead of implying it's done.

Context & auto-compact

The status bar shows ctx 4.1k/8k — the size of the last prompt vs. the model's context window. The window is auto-detected from LM Studio's /api/v0/models; set llm.context_window to override (0 disables auto-compact). When the prompt passes ~75% of the window the session is auto-compacted into a short summary to free room (run it any time with /compact; the threshold and whether it summarizes at all are configurable — see memory/compact_threshold below). Every task's goal and outcome is also archived in full, never summarized, to sessions/<name>.archive.jsonl in the workspace's state directory — once there's something in it, the model gets a history tool to recall or search past tasks that a compaction summary has since shortened.

REPL commands

Anything not starting with / is run as a task. Tab completes commands.

command what
/plan, /auto plan mode (propose only) vs auto mode (execute) — also shift+tab
/goal <text> set & pursue a multi-turn goal (judge re-feeds until met); go · clear · ttl <n> · off
/skills list / toggle / install on-demand instruction packs
/permissions relax approval for non-destructive file/shell actions
/status config, knowledge base, and trace paths
/budget [usd] cap estimated spend per run — refuses new tasks once hit; off disables
/diff show uncommitted workspace changes (what the agent changed), colorized
/steer <note> fold a note into a running task without stopping it — lands on its next step (esc cancels instead)
/btw <question> ask a quick side question mid-task — one-turn answer, no tools, the task keeps going
/snip [name] prompt templates — /snip <name> pulls a saved template into the input to edit & send; save <name> [text] (omit text → your last prompt) · list · rm <name>
/usage token spend + estimated $ (today / 7d / 30d / all, by day, by model); clear · purge <days> · retain <days>
/login (re)configure server URL / model / key, then reload
/new [name] start a NEW session (the old one stays in /sessions)
/clear wipe this session's context + the screen (same session)
/compact summarize the session so far to free up context
/color [name] change the TUI frame color (cycles if no name)
/rename <name> rename the agent (saved in settings)
/sessions pick a saved session to switch to (/sessions <name> switches, /sessions delete <name> deletes)
/agents sub-agent profiles: add (LLM) / add-tool (external CLI) / rm / exec
/telemetry anonymous usage statistics: exactly what is sent · on · off · reset
/rate rate ipsupport-code: /rate <1-5> <a few words> [--name <you>] · later · never
/loop <interval> [xN] <task> re-run a task on an interval (e.g. /loop 5m <task>, /loop 30s x10 <task>); esc stops it
/help command list
/exit, /quit leave

The input is multi-line: paste a whole block (e.g. a YAML snippet) and it keeps its line breaks, the box grows and word-wraps instead of scrolling on one line, and alt+enter (or ctrl+j) inserts a newline by hand. Enter submits.

History. With an empty input, ↑ / ↓ recall previous messages to re-run or fix a typo — the first ↑ jumps to your last prompt. History is persisted per workspace (history in its state directory), so recall spans past runs; /history lists recent prompts and /history <text> filters them, and ctrl+r opens an incremental reverse-search (type to narrow · ctrl+r older · enter use). Tab completes /commands and @file paths against the workspace. (PgUp/PgDn and the wheel scroll the log.)

Everything you type is a message queue. While a task runs the input stays live: Enter queues the next message — a task or a /command — pinned above the input and drained in order when the task finishes (deferred commands are no longer dropped). ↑ on an empty input pulls the last queued message back to edit or drop, and esc cancels.

Approvals. When the agent asks to approve a file write or shell command, it takes over the keys: press y (approve), n (deny) or a — allow every action of that kind (file/shell/git/spawn/external) for the rest of the session; /permissions shows what you allowed and /permissions reset revokes it. ↑/↓ toggle the explicit Yes/No prompt; esc backs out to keep typing (queue a message first) without answering — the approval just keeps waiting.

Shell. /shell (or !) drops you into an interactive shell in the workspace — do things by hand, exit to return. !cmd runs a single command and shows its output. These are your commands, not the agent's, so they aren't gated by the permission policy.

Custom system prompt. The built-in engine prompt is deliberately tiny; you can replace it with .agent/system.md (per project) or ~/.config/ipsupport-code/system.md (global). ipsupport-code -dump-prompt prints the default to start from (> .agent/system.md). Your CLAUDE.md, environment, and skills are still appended after it. /status shows which prompt is in effect. (A bloated prompt makes a small model call tools worse — edit at your own risk.)

How it works

  • Native tool calling. Talks to any OpenAI-compatible server — LLMTray, LM Studio, Ollama, vLLM, a LiteLLM proxy or a cloud provider — and lets the model call tools natively. One client for all of them; llm.base_url / llm.api_key or a providers entry picks the endpoint.
  • Fat tools. One tool per domain, each {"action": ..., "params": {...}}. The catalog stays tiny (~1k tokens) so small models prefill fast and route well; a declarative Domain generates each tool's schema, help, and validation.
  • Proactive help. When a tool fails, a matching lesson from past runs is injected straight into the error the model sees — it doesn't have to ask.
  • Reflection. After a task, a second model pass distills durable lessons into ~/.config/ipsupport-code/knowledge.json (env-general tool pitfalls) and durable facts about the current project (build/test/run commands, where things live, conventions) into facts.json in the workspace's state directory — folded into the prompt next run. Each lesson tracks when it was last seen (bumped on recurrence); /knowledge reports the store and clear / purge <days> / retain <days> prune stale ones (retain auto-purges on startup) so the memory doesn't accrete junk forever.
  • Code search. The file tool's search action greps the workspace by regex (file:line: match), skipping VCS/dep/build dirs and binaries — no external grep.
  • Session memory. Remembers your goals and its answers across turns and across restarts, kept per workspace and per agent name (sessions/<name>.json in the workspace's state directory). On startup the TUI shows a navigable chooser of saved sessions (↑↓ / enter / d) — and on restore it replays the recent exchanges so you pick up where you left off. /new <name> starts a fresh named thread; /new wipes the active one.
  • Cache-friendly prompts. A local server reuses its prompt cache only while each request starts with exactly what the last one did, so within a session the prompt only grows at the end. The system prompt stays as the session began: facts learned after a task, and the plan-mode directive, arrive as an <agent-note> on the next task's message instead of rewriting it. On a local provider the goal judge and the reflection pass go out as the next message of the task itself — same prompt, same tools — so they are answered from the cache instead of prefilling the whole record again (they fall back to their own prompt when that yields no answer).
  • State outside the project. What the agent writes for itself — goal, facts, lessons, prompt history, sessions — lives in ~/.config/ipsupport-code/state/<workspace>-<hash>/, never in the workspace, so a model listing the project can't find and replay its own state. Files you write (.agent/config.json, system.md, judge.md, compact.md, instructions.md) stay in the project. Older installs are moved out of .agent/ once, on first start.
  • Resilience. Exponential-backoff retry on transient 5xx/network errors, an idle watchdog that aborts a silently-stalled stream, and a stuck-loop guard.
  • Project instructions. Reads a CLAUDE.md / AGENTS.md / .agent/instructions.md from the workspace into the system prompt.
  • Trace = dataset. Every step (goal, tool call, observation, final, lesson) is appended as JSONL to ~/.config/ipsupport-code/traces.jsonl.
  • Usage ledger. Token spend is recorded per day and per provider/model to ~/.config/ipsupport-code/usage.json and accumulates across runs; /usage shows today / 7-day / 30-day / all-time rollups, with clear, purge <days>, and a retain <days> retention window.

Configuration

Settings merge over safe defaults from two JSON files:

  • ~/.config/ipsupport-code/config.json — machine-level: the llm connection (server URL, model, key, context_window). Written by first-run setup.
  • <workspace>/.agent/config.json — per-project: the permission policy (see .agent/config.example.json). Wins over the user file for everything EXCEPT the llm connection and providers — a workspace is a checkout you might not fully trust, so it can tighten or loosen what the agent may do, but it can't redirect your model endpoint or add a provider preset while your real API key still gets sent wherever it points.

Edit either file from a script with the config subcommand — no interactive session needed:

ipsupport-code config set update_check false   # any key, dotted paths for nested
ipsupport-code config set goal_max_returns 8   # JSON values keep their type
ipsupport-code config get llm.model            # print the effective value
ipsupport-code config unset channel            # back to the default
ipsupport-code config list                     # effective config, one key per line
ipsupport-code config --local set run.default allow   # write the workspace file

Values that parse as JSON keep their type (false, 8, ["x"]); anything else is a literal string. set writes the global user file by default (--local targets <workspace>/.agent/config.json); a wrong type or a misspelled key is rejected rather than saved. get/list show the effective, merged value — a key sitting at its zero default (e.g. offline when off) prints nothing, like git config --get. The workspace file always wins over the global one — e.g. the interactive /config panel's file/run permission rows persist to the workspace file (they're a per-project concern), so setting the same key globally afterward has no visible effect; set/unset warn on stderr when that's about to bite you.

run.timeout_seconds caps how long a shell command may run (default 60s); raise it for slow builds/test suites, or let the model pass a larger per-call timeout.

On Windows, run.shell picks the shell commands run in: unset uses pwsh if installed, else Windows PowerShell; or set pwsh, powershell or cmd (ipsupport-code config set run.shell cmd). Elsewhere it is always sh.

Cross-task memory: once the context window fills past compact_threshold (default 0.75), the session is folded into an LLM-written recap to free headroom — memory raw turns that off entirely, keeping turns verbatim and only dropping the oldest ones outright (no paraphrase) once there are too many (500 messages — high enough that it's a rare backstop, not routine): ipsupport-code config set memory raw / config set compact_threshold 0.85, or toggle/cycle both live from the /config panel. Either kind of cut — an LLM recap or the plain drop — changes the front of the prompt, which breaks a local server's KV-cache reuse for it (matched only on an identical prefix) just as much as the other; raw's high cap exists so that cost stays rare instead of hitting on every turn once a session runs long.

Permissions for run and file resolve per action: a deny glob blocks, an allow glob runs without asking, otherwise the default (ask/allow/deny) applies. Run-command deny globs match anywhere in the command (so rm -rf* catches cd x && rm -rf /); file globs are path-aware (**, *.go) and confined to jail. A run allow glob must match each chained command whole, and * means any characters: ls* also allows lsof -i. To allow just ls with arguments, write ls * (plus ls for the bare command).

The protective deny floor is always unioned in — your config adds to it, it can't remove it. It judges the words the shell will actually run — after quotes and backslashes are removed, through command/exec/env/xargs and similar wrappers, across ;, &&, ||, | and & — and the deny and secret-file floors ignore case, since macOS and Windows file systems do. It is still a static check: sh -c "…", variables and encoded commands get past it, and the ask default is what stops those.

Logging: IPS_LOG=debug|info|warn|error (default warn). In the TUI it writes to ~/.config/ipsupport-code/agent.log (so raw log lines don't bleed over the screen) — tail -f it to watch retries/warnings live; a piped/one-shot run logs to stderr.

Layout

cmd/agent          CLI, plain REPL, the Bubble Tea TUI, external CLI-agent runner
internal/llm        OpenAI-compatible client (streaming, retry, context detection)
internal/agent      the reason → act → observe loop (+ plan mode, goal judge)
internal/tool       fat tools: file, run, git, web, calc, agent, mcp, skill, help, history
internal/skill      downloadable, toggleable instruction packs
internal/policy     workspace permission engine (+ jail, deny floor)
internal/knowledge  persistent pitfall store
internal/reflect    post-task lesson + project-fact distillation
internal/trace      JSONL decision trace (the dataset)
internal/config     config load/merge
internal/mcp        MCP client (stdio + HTTP)
internal/usage      token/cost ledger (powers /usage and /budget)
internal/selfupdate checksum-verified in-place self-update
internal/risk       risk scorer behind the risk check (pure-Go inference, embedded weights)
internal/sandbox    opt-in OS sandbox for `run` (Seatbelt on macOS, Landlock on Linux)
internal/procgroup  kill a command's whole process tree on cancel/timeout
internal/filelock   cross-process lock for files shared between running sessions
internal/e2e        end-to-end tests: real loop, tools and policy against a fake server
internal/textutil · internal/atomicfile   shared helpers (clipping, atomic writes)

Contributing

See CONTRIBUTING.md. CI runs gofmt, go vet, the race suite, and a cross-compile of every target on each push and PR.

BACKLOG.md lists work that is designed but not built — each entry says what is wrong, what the fix is, and what that fix would break.

adr/ records the architecture decisions the code rests on — where state lives, what a workspace may override, why external agents are gated separately, how the risk model is trained and shipped.

License

MIT © ipsupport-llc

The binary carries the licenses of everything it links (all MIT or BSD) and of the data its risk model was trained from: /license lists them, ipsupport-code --license prints the full texts. After changing dependencies, regenerate them with go run ./internal/legal/gen; a test fails until you do.

About

A small self-learning coding agent for your own model — any OpenAI-compatible server, local (LLMTray, LM Studio, Ollama, vLLM) or cloud. Single static Go binary, plan/auto modes, skills, permissions, Claude-Code-style TUI.

Topics

Resources

Contributing

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages