A sandboxed JS-syntax scripting language that gives an LLM one eval tool
instead of twenty. It earns that one tool on three things a pile of discrete
tools can't offer at once:
- Reliability on LLM-hard tasks — exact letter counting, arithmetic, string and data wrangling: right by construction and faster than reasoning it out.
- ~96% less tool-schema context — one
eval+ a compact reference (or a deferredhelp()) replaces one tool schema per capability, and stays flat as you add capabilities (KV-cache friendly). - Safe untrusted execution — no imports,
eval,process, prototype chain, or ambient I/O, plus step/time limits, so model-authored code can't reach or hang the host.
Measured against the reasoning-only baseline — 96% context cut, 3.3× faster on
LLM-hard tasks, every host escape blocked — in benchmarks/.
Status: complete. Parser, runtime, interpreter, all toolkits (core, math, web, file, graph, SQL, browser), the MCP server, and the LLM benchmark harness are implemented and tested — 750+ tests.
go install github.com/iodesystems/mcpshell/cmd/mcpshell@latest
This drops a mcpshell binary in your $GOBIN. Or run it without installing:
go run github.com/iodesystems/mcpshell/cmd/mcpshell@latest '[1,2,3] |> map(x => x * 10)'
Pin a release with @v0.1.0 in place of @latest. Requires Go 1.26+ — the
go directive auto-fetches the toolchain for anyone on Go 1.21 or newer. No
Java is needed; the generated parser is committed.
./bin/mcpshell '[1,2,3] |> map(x => x * 10)' # evaluate code
echo 'range(5) |> map(n => n * n)' | ./bin/mcpshell
./bin/mcpshell # interactive REPL
./bin/mcpshell --prompt # print the LLM system prompt
./bin/mcpshell mcp --web --files-dir ./data # run as an MCP server over stdio
As an MCP server, mcpshell exposes eval, help, and prompt tools. It can
also compose upstream MCP servers as namespaced commands
(--connect/--mcp); an upstream that is itself mcpshell is skipped to avoid
a recursive eval loop.
Numbers are exact and arbitrary-precision — no silent float64 rounding.
Integer arithmetic auto-promotes past 2^53, decimals and rationals stay exact,
and only transcendental ops (sqrt/sin/log/non-integer powers) are
floating-point (the precision boundary). A tiered int64 → big.Int → big.Rat
representation keeps the common integer path allocation-free (~1.3× the float
baseline on hot loops).
2 ** 1000 // 302 exact digits
0.1 + 0.2 == 0.3 // true
1/3 + 1/3 + 1/3 // 1
Every eval is bounded (step, call-depth, wall-clock, and output limits) so
untrusted code can't hang the host. profile(() => …) returns
{result, steps, lines} to see which lines cost the most, and a tripped step
limit names the hot lines instead of just a number.
Toolkits register namespaced commands into the shell. core, math, and
graph are always on; the rest are opt-in via mcpshell mcp flags so the
LLM only sees the capabilities you grant.
| Toolkit | Enable with | Commands |
|---|---|---|
| core | always on | strings, arrays, objects, JSON, regex, pipes, control flow |
| math | always on | arithmetic, rounding, trig, min/max/sum, etc. |
| graph | always on | in-memory node/edge graph: addNode, link, nodes, setProps, … |
| web | --web |
Web.fetch, Web.fetchText, Web.search, Web.clearCache; Html.select, Html.links, Html.text, Html.table |
| file | --files-dir DIR |
sandboxed reads/writes rooted at DIR (add --files-read-only) |
| sql | --sql 'ns=DSN' |
Postgres + SQLite: ns.query, ns.tables, ns.columns, ns.schema, ns.execute |
| browser | --browser |
headless-Chrome automation (chromedp): Browser.open, click, type, text, html, select, wait, screenshot, eval |
Attach one or more databases; each becomes a namespace. The DSN is a SQLite
path or a postgres:// URL. Queries are read-only by default (only
select/with/show/… are allowed) and capped at 500 rows; pass
--sql-writable to permit writes and DDL via ns.execute.
mcpshell mcp --sql 'app=postgres://user:pass@localhost/app' \
--sql 'cache=./local.sqlite'
# then, in eval: app.query("select id, email from users where id = $1", [42])
# app.schema() · cache.tables("order")
--browser drives a real headless Chrome through chromedp — navigate, click,
fill inputs, extract text/HTML, run in-page JS, and screenshot. Requires a
Chrome/Chromium install on the host.
mcpshell mcp --browser
# Browser.open("https://example.com") · Browser.text("h1")
# Browser.select("table tr") · Browser.screenshot("shot.png")
bin/mcpshell is a launcher: it builds a per-arch binary (bin/mcpshell-<os>-<arch>)
on demand — when missing, when any source is newer, or when BUILD=true is set —
then execs it. The same launcher works across machines and architectures.
The three claims above are measured in benchmarks/
(context cost, sandbox properties, and a with/without-tool showcase per suite).
bin/bench runs the challenge suite against an OpenAI-compatible endpoint and
can run the same problems without the tool to measure the difference (scored on
correctness, turns, processed/cached tokens, time):
./bin/bench # full suite
./bin/bench --only llm_hard_ # one category
./bin/bench --no-tool # reasoning-only baseline (no tool)
./bin/bench compare with_dir without_dir # side-by-side comparison doc
./bin/bench context # tool-context cost measurement
Results are written as markdown + a machine-readable results.json. Every
deterministic teaser also has a verified single-eval reference solution
(bench/references.go, checked by go test ./bench/ with no LLM). bin/bench
auto-builds the same way bin/mcpshell does.
grammar/ ANTLR4 grammar (.g4)
parser/ Generated Go parser + hand-written lexer base (committed)
runtime/ Value types, environment, interpreter, the Shell facade
toolkit/ Built-in command toolkits — core, math, web, file, graph, sql, browser
mcp/ MCP server + upstream-MCP client/composition
bench/ LLM benchmark harness (OpenAI-compatible agent loop)
cmd/ Binaries: cmd/mcpshell (the CLI), cmd/bench (the benchmark)
bin/ Auto-building launchers — bin/mcpshell, bin/bench
make build # go build ./...
make test # go test ./...
make generate # regenerate parser/ after editing grammar/ (needs Java + ANTLR 4.13.2)
Requires Go 1.26+. Parser regeneration requires a JDK and the ANTLR 4.13.2 tool
jar with its dependencies (resolved from the local Maven repo — see Makefile).