Named for Voltaire's Pangloss, who insisted this is the best of all possible worlds.
I built Pangloss to test whether a fusion of diverse coding agents beats the best single model, then I measured it. On a 9-task SWE-bench Lite subset with a hidden grader, fusion scored the same as the best solo lane, 1 of 9. Every lane solved the same one task and failed the same eight, so even a perfect selector could not have done better. Cross-model review scores did not track hidden-test correctness: the only lane that solved anything won 0 of 9 review votes. An earlier polyglot run showed fusion ahead only when the grading tests were visible in the worktree. The lever is a trustworthy validation signal, not model diversity.
Full method, per-task analysis, and caveats: bench/swe/FINDINGS.md.
Pangloss runs a diverse roster of AI coding agents in parallel, with different models and different agentic harnesses, has them cross-review each other's work, and synthesizes one result. The idea came from OpenRouter's Fusion: a loop of sonnet-level and open-weight models might rival any single frontier model, and in combination might exceed it. This repo is the apparatus I used to test that, and the Result section above is what it measured.
Each run is a loop over four phases. Every agent works in its own git worktree, an isolated checkout on its own branch. Agents run on the host, so each tool uses its own existing auth.
feature request
|
v PLAN N models draft independently. A rotating synthesizer merges
| them into one canonical plan, with an optional approval gate.
v CODE Each model implements the plan in its own worktree, runs the
| project's build and tests, iterates to green, commits.
v REVIEW Every model reviews every implementation: score, novel ideas,
| gaps, what is still needed.
v SELECT A weighted vote (self-reviews down-weighted) picks a winner and
| produces a revision brief.
+--> If not converged: re-base every agent on the winning branch, turn the
brief into a revision plan, and loop. Stop on convergence or round cap.
Even when several agents pass, the review phase surfaces each one's distinct ideas, and the revision rounds graft them into the kept winner.
Agents are <tool> + <model> pairs. The same model through two different
harnesses, say Sonnet via claude-code and via codex to OpenRouter, produces
different work.
| Tool | How it's driven | Examples |
|---|---|---|
| claude | Claude Code CLI (claude -p) |
claude:sonnet, claude:opus |
| codex | OpenAI Codex CLI (codex exec) |
codex:gpt-5 |
codex --oss |
local open-weight via Ollama | oss:gpt-oss:120b |
| cursor | Cursor Agent CLI (cursor-agent -p) |
cursor:claude-4.5-sonnet, cursor:kimi-k2.5 |
| OpenRouter | any OpenRouter model via codex's custom provider | openrouter:qwen/qwen3-coder, openrouter:z-ai/glm-4.6 |
| gemini | Gemini CLI (gemini -p) |
gemini:gemini-2.5-pro (needs GOOGLE_CLOUD_PROJECT) |
Drop any of these into a roster ad hoc (--roster "openrouter:qwen/qwen3-coder,claude:sonnet,oss:gpt-oss:120b"),
or use a named roster from pangloss.config.json (open-weight-heavy,
frontier, sonnet-family, openrouter).
npm ci
npm run build
# See the roster catalog and named rosters
node dist/cli.js agents
# Check that your roster's CLIs are installed and authed, and preview invocations
node dist/cli.js doctor --roster open-weight-heavy
# Run on the current repo (interactive: asks for the change, approves the plan)
node dist/cli.js run
# Or fully unattended
node dist/cli.js run --non-interactive --yes \
--roster "claude:sonnet,openrouter:qwen/qwen3-coder,oss:gpt-oss:120b" \
--request "Add a slugify() utility with unit tests"Artifacts for every run land under .pangloss/runs/<run-id>/, with a per-round
plan.json, code-outcomes.json, reviews.json, selection.json, and
summary.md. The winning worktree is kept for inspection.
| Command | Description |
|---|---|
pangloss run |
Plan, code, review, select, then the revise loop. The default command. |
pangloss agents |
List configured rosters and agent presets. |
pangloss doctor [--roster ...] |
Verify roster CLIs are installed and authed, preview invocations. |
pangloss models [--filter ...] |
List OpenRouter model slugs usable as openrouter:<slug>. |
pangloss config / pangloss setup |
Generate pangloss.config.json or .env. |
Key run flags: --roster <name|csv>, --request <text>, --rounds <n>,
--keep-worktrees, --timeout <min>, -y/--yes, --non-interactive.
Requirements: Node 20 or newer, Git 2.5 or newer, and the CLIs for whatever
roster you use (claude, codex, cursor-agent, gemini, ollama). Each tool
authenticates itself, by subscription OAuth or API key. Pangloss runs them on the
host, so no keys are injected anywhere.
OpenRouter is optional, for the openrouter: lanes and the openrouter roster.
Set OPENROUTER_API_KEY in .env. Discover model slugs with pangloss models.
Local open-weight lanes (oss:) need ollama running with the model pulled, for
example ollama pull gpt-oss:120b. They are driven via codex exec --oss.
The behavior contract every model obeys inside its worktree lives in
.claude/skills/pangloss-worktree/SKILL.md
and is injected into every agent's system prompt.
For targets that need a database or other services, the compose runtime gives
every agent its own stack, so they run in parallel without host-port collisions
and your source tree is never modified. Add a compose block to the manifest
pointing at the app's existing docker-compose.yml:
Per agent, Pangloss writes a port-rewritten copy of your compose file, leaving
the source untouched, brings it up under a unique project name
(docker compose -p pangloss-<run>-<agent>), waits for the database, runs
dbSetup, injects the connection URL, then builds, tests, and runs e2e against
that isolated stack before tearing it down with down -v. Agents still run on
the host, so their CLI auth just works. Only the runtime is containerized.
If the target's tests pin a database port, that guard will reject the remapped
port. Widen the guard to accept any port, keep its credential and database-name
check, and mark the file skip-worktree so no lane commits the change.
Run it from a clone you have dedicated to Pangloss. It makes its own worktrees and branches there, so your other clones stay free for manual work.
cd ~/projects/your-app
node /path/to/pangloss/dist/cli.js run \
--non-interactive --yes --overnight \
-c /path/to/pangloss/examples/web-app-postgres.config.json \
--roster "claude:sonnet,oss:gpt-oss:120b,openrouter:qwen/qwen3-coder" \
--request "...your feature..."Use --overnight here. Heavy e2e runs and slow local models want no clock. See
examples/web-app-postgres.config.json.
This path was validated end to end on a private Next.js app, including an isolated per-lane Postgres as the selection gate.
npm run build # tsc
npm test # jest
npm run lint # eslint
npm run typecheck # tsc --noEmitMIT. See LICENSE.