Skip to content
brookrPublic

About

Parallel, multiplexing LLM code agent system

Resources

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

Pangloss

CI

Named for Voltaire's Pangloss, who insisted this is the best of all possible worlds.

Result

I built Pangloss to test whether a fusion of diverse coding agents beats the best single model, then I measured it. On a 9-task SWE-bench Lite subset with a hidden grader, fusion scored the same as the best solo lane, 1 of 9. Every lane solved the same one task and failed the same eight, so even a perfect selector could not have done better. Cross-model review scores did not track hidden-test correctness: the only lane that solved anything won 0 of 9 review votes. An earlier polyglot run showed fusion ahead only when the grading tests were visible in the worktree. The lever is a trustworthy validation signal, not model diversity.

Full method, per-task analysis, and caveats: bench/swe/FINDINGS.md.

The hypothesis

Pangloss runs a diverse roster of AI coding agents in parallel, with different models and different agentic harnesses, has them cross-review each other's work, and synthesizes one result. The idea came from OpenRouter's Fusion: a loop of sonnet-level and open-weight models might rival any single frontier model, and in combination might exceed it. This repo is the apparatus I used to test that, and the Result section above is what it measured.

How it works

Each run is a loop over four phases. Every agent works in its own git worktree, an isolated checkout on its own branch. Agents run on the host, so each tool uses its own existing auth.

feature request
   |
   v  PLAN     N models draft independently. A rotating synthesizer merges
   |           them into one canonical plan, with an optional approval gate.
   v  CODE     Each model implements the plan in its own worktree, runs the
   |           project's build and tests, iterates to green, commits.
   v  REVIEW   Every model reviews every implementation: score, novel ideas,
   |           gaps, what is still needed.
   v  SELECT   A weighted vote (self-reviews down-weighted) picks a winner and
   |           produces a revision brief.
   +--> If not converged: re-base every agent on the winning branch, turn the
        brief into a revision plan, and loop. Stop on convergence or round cap.

Even when several agents pass, the review phase surfaces each one's distinct ideas, and the revision rounds graft them into the kept winner.

The roster

Agents are <tool> + <model> pairs. The same model through two different harnesses, say Sonnet via claude-code and via codex to OpenRouter, produces different work.

Tool How it's driven Examples
claude Claude Code CLI (claude -p) claude:sonnet, claude:opus
codex OpenAI Codex CLI (codex exec) codex:gpt-5
codex --oss local open-weight via Ollama oss:gpt-oss:120b
cursor Cursor Agent CLI (cursor-agent -p) cursor:claude-4.5-sonnet, cursor:kimi-k2.5
OpenRouter any OpenRouter model via codex's custom provider openrouter:qwen/qwen3-coder, openrouter:z-ai/glm-4.6
gemini Gemini CLI (gemini -p) gemini:gemini-2.5-pro (needs GOOGLE_CLOUD_PROJECT)

Drop any of these into a roster ad hoc (--roster "openrouter:qwen/qwen3-coder,claude:sonnet,oss:gpt-oss:120b"), or use a named roster from pangloss.config.json (open-weight-heavy, frontier, sonnet-family, openrouter).

Quick start

npm ci
npm run build

# See the roster catalog and named rosters
node dist/cli.js agents

# Check that your roster's CLIs are installed and authed, and preview invocations
node dist/cli.js doctor --roster open-weight-heavy

# Run on the current repo (interactive: asks for the change, approves the plan)
node dist/cli.js run

# Or fully unattended
node dist/cli.js run --non-interactive --yes \
  --roster "claude:sonnet,openrouter:qwen/qwen3-coder,oss:gpt-oss:120b" \
  --request "Add a slugify() utility with unit tests"

Artifacts for every run land under .pangloss/runs/<run-id>/, with a per-round plan.json, code-outcomes.json, reviews.json, selection.json, and summary.md. The winning worktree is kept for inspection.

CLI

Command Description
pangloss run Plan, code, review, select, then the revise loop. The default command.
pangloss agents List configured rosters and agent presets.
pangloss doctor [--roster ...] Verify roster CLIs are installed and authed, preview invocations.
pangloss models [--filter ...] List OpenRouter model slugs usable as openrouter:<slug>.
pangloss config / pangloss setup Generate pangloss.config.json or .env.

Key run flags: --roster <name|csv>, --request <text>, --rounds <n>, --keep-worktrees, --timeout <min>, -y/--yes, --non-interactive.

Setup

Requirements: Node 20 or newer, Git 2.5 or newer, and the CLIs for whatever roster you use (claude, codex, cursor-agent, gemini, ollama). Each tool authenticates itself, by subscription OAuth or API key. Pangloss runs them on the host, so no keys are injected anywhere.

OpenRouter is optional, for the openrouter: lanes and the openrouter roster. Set OPENROUTER_API_KEY in .env. Discover model slugs with pangloss models.

Local open-weight lanes (oss:) need ollama running with the model pulled, for example ollama pull gpt-oss:120b. They are driven via codex exec --oss.

The behavior contract every model obeys inside its worktree lives in .claude/skills/pangloss-worktree/SKILL.md and is injected into every agent's system prompt.

Web apps: an isolated database per agent

For targets that need a database or other services, the compose runtime gives every agent its own stack, so they run in parallel without host-port collisions and your source tree is never modified. Add a compose block to the manifest pointing at the app's existing docker-compose.yml:

"manifest": {
  "setup": "npm ci",
  "build": "npm run build",
  "test":  "npm test",              // integration tests hit the DB
  "e2e":   "npm run test:e2e:ci",   // optional; Playwright self-starts the app
  "compose": {
    "file": "docker-compose.yml",
    "dbService": "db",
    "dbPortBase": 5440,             // agent i gets 5440 + i
    "urlEnv": "DATABASE_URL",
    "urlTemplate": "postgres://test_user:test_password@localhost:{port}/test_db",
    "dbSetup": "npm run db:rebuild" // migrate and seed the fresh DB
  }
}

Per agent, Pangloss writes a port-rewritten copy of your compose file, leaving the source untouched, brings it up under a unique project name (docker compose -p pangloss-<run>-<agent>), waits for the database, runs dbSetup, injects the connection URL, then builds, tests, and runs e2e against that isolated stack before tearing it down with down -v. Agents still run on the host, so their CLI auth just works. Only the runtime is containerized.

If the target's tests pin a database port, that guard will reject the remapped port. Widen the guard to accept any port, keep its credential and database-name check, and mark the file skip-worktree so no lane commits the change.

Run it from a clone you have dedicated to Pangloss. It makes its own worktrees and branches there, so your other clones stay free for manual work.

cd ~/projects/your-app
node /path/to/pangloss/dist/cli.js run \
  --non-interactive --yes --overnight \
  -c /path/to/pangloss/examples/web-app-postgres.config.json \
  --roster "claude:sonnet,oss:gpt-oss:120b,openrouter:qwen/qwen3-coder" \
  --request "...your feature..."

Use --overnight here. Heavy e2e runs and slow local models want no clock. See examples/web-app-postgres.config.json.

This path was validated end to end on a private Next.js app, including an isolated per-lane Postgres as the selection gate.

Development

npm run build       # tsc
npm test            # jest
npm run lint        # eslint
npm run typecheck   # tsc --noEmit

License

MIT. See LICENSE.

About

Parallel, multiplexing LLM code agent system

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages