MCP server that lets any agent or MCP host delegate to headless Codex, Grok, Claude Code, and OpenCode CLIs, with Cursor available as an opt-in fallback. Use the fleet for implementation, planning, and project exploration without burning the caller's context on raw worker output.
Worker tools take optional model and effort overrides, return a session_id, and support
follow_up. Difficulty levels select a distinct default model across the three active
subscriptions; delegate also accepts an explicit engine, including pay-per-token OpenCode.
The server exposes twelve tools:
| Tool | Purpose |
|---|---|
delegate |
Run a task with full read/edit/shell access in cwd. Required level: 1=GPT-6 Luna max (codex), 2=GPT-6 Sol high (codex), 3=GPT-6 Sol max (codex), 4=Claude Opus 5.5 high (claude), 5=Claude Opus 5.5 max (claude). Levels 4 and 5 are expensive — 5 costs ~3.3× level 4; last resort only. Both share the Claude Code host subscription. Optional engine overrides the tier; opencode requires a provider/model model. Optionally accepts an agent persona by name or inline {prompt}. |
fast_delegate |
Same full read/edit/shell access as delegate, but with no level to pick: it routes to whichever CLI is currently the fastest and healthy. Prefer it over delegate when the task is simple or urgent and picking a level is not worth it. First two candidates are subscriptions (GPT-6 Luna medium on codex, then Claude Haiku low); pay-per-token OpenRouter (mercury-2) is the 3rd fallback, only after codex and claude are missing, quota-exhausted or unhealthy. The accepted Claude-subscription cost is the same one used by a Claude Code host orchestrator. Optionally accepts an agent persona. |
explore |
Read-only exploration on Codex with gpt-6-luna and explicit medium effort by default (POLYAGENT_EXPLORE_EFFORT). question alone → broad fan-out search returning file:line refs; question+files → answer about those files; neither → general project map. breadth: "thorough" sweeps wider. Locates, does not review. |
read_slice |
Surgical read-only read: returns ONLY the code relevant to want (exact lines with file:line) from the given files — the full file never enters your context. Codex uses gpt-6-luna with explicit medium effort by default (POLYAGENT_EXPLORE_EFFORT). Use instead of reading large files whole. |
run_filtered |
Run a shell command with full access and get back ONLY the lines relevant to want — semantic filtering of huge build/test/log output. Default engine is the same FAST_CANDIDATES cascade as fast_delegate: codex GPT-6 Luna medium, then Claude Haiku, then OpenRouter mercury-2. OpenRouter credit is spent only after both subscription candidates are missing, quota-exhausted or unhealthy. |
web_lookup |
Web/docs lookup through Codex/GPT-6 Luna with explicit medium effort by default (POLYAGENT_EXPLORE_EFFORT), real web search enabled and a read-only filesystem. Falls back to Claude Haiku (native WebSearch) when codex is unhealthy or out of quota. |
decide |
Ask TypeSafe's Jev model for calibrated probabilities or a typed choice label. Pay-per-token through OpenRouter; useful for risky-call gates, classification, and verifying worker claims. |
generate_image |
Generate or edit an image through Codex's built-in image tool and save it inside cwd. |
fan_out |
Get independent opinions, compare approaches, cross-check a risky verdict, or split broad research. Avoid simple lookups, single-file edits, and tightly coupled sequential work; it runs several workers and costs several times one delegate. With the optional Jev consensus gate, high agreement skips the Codex arbiter; uncertainty or Jev failure keeps it. |
follow_up |
Continue a prior session by session_id. |
bridge_stats |
Report calls and chars returned to context per tool, plus ratings; optional export: true writes research/bench/<YYYY-MM-DD>-ratings.md (needs POLYAGENT_LOG). |
rate |
Grade a reviewed result from 1–5 by its session_id; ratings stay local and feed bridge_stats. |
Worker tools accept cwd, model, and effort where applicable. delegate requires a level
(1-5); fast_delegate has none and picks the fastest healthy engine. Explicit model/effort
values override the selected tier.
A call that fails because the engine's plan quota is exhausted does not silently retry on
another engine — spending the next subscription is your decision. The call fails with an actionable
error naming the engines still available (installed, enabled, and capable of what that tool needs)
and how to switch: engine:"<x>" on the four auxiliary tools and delegate, or the lowest
still-usable level:<n> when a tier engine remains. Tools that pick the engine themselves (fast_delegate, fan_out) and follow_up
(pinned to the resumed session's engine) report the quota without suggesting a parameter, and
generate_image reports it against the two engines that have an image tool at all (codex, grok). A transient rate limit is reported separately and asks you to wait, since switching
engines would not help. Anything the classifier does not recognize — an expired login, for one —
propagates as the raw CLI failure instead of being guessed at.
- Node ≥ 18
bubblewrap(bwrap) installed — required, not recommended: the sandbox is mandatory and the server refuses to start without it (sudo apt install bubblewrap). OnlyPOLYAGENT_SANDBOX=offwaives it, as an explicit operator choice.- Codex installed and authenticated for read tools and levels 1/3; Grok for levels 2/4; Claude Code for level 5.
- Optional Cursor fallback: install
cursor-agentand setPOLYAGENT_ENABLE_CURSOR=1.
Installing via an AI agent? Point it at
INSTALL.md— an agent-facing, copy-paste guide that detects the host and registers the bridge in Claude Code, Cursor, Codex, Grok, or any generic MCP host.
git clone https://github.com/JaimeJunr/polyagent-mcp.git
cd polyagent-mcp
npm install
npm run buildClaude Code:
claude mcp add polyagent -s user -- node /abs/path/to/polyagent-mcp/dist/index.jsAny host — add to its mcp.json:
{
"mcpServers": {
"polyagent": {
"command": "node",
"args": ["/abs/path/to/polyagent-mcp/dist/index.js"]
}
}
}Permissions (Claude Code): claude mcp add registers the server but does not grant
tool permission — without an allowlist every bridge call prompts for approval. After
registering, add either "mcp__polyagent__*" (full; also auto-approves mutating tools
delegate/fast_delegate/run_filtered/follow_up) or a read-only subset
(explore/read_slice/web_lookup/bridge_stats) under
permissions.allow in settings.json. Full options and trade-offs:
INSTALL.md §3. Cursor/Codex/other hosts
have their own approval settings — consult the host.
| Var | Default | Meaning |
|---|---|---|
POLYAGENT_CURSOR_BIN |
cursor-agent |
Path to the optional Cursor CLI fallback. |
POLYAGENT_GROK_BIN |
grok |
Path to the Grok CLI. |
POLYAGENT_CODEX_BIN |
codex |
Path to the Codex CLI. |
POLYAGENT_CLAUDE_BIN |
claude |
Path to the Claude Code CLI. |
POLYAGENT_OPENCODE_BIN |
opencode |
Path to the OpenCode CLI used by explicit engine selection. |
POLYAGENT_ENABLE_CURSOR |
(off) | Set to 1/true to allow Cursor fallback when a tier's preferred CLI is missing. Otherwise the call fails with the missing CLI named. |
POLYAGENT_MODEL |
composer-2.5-fast |
Default model for the optional Cursor path. |
POLYAGENT_EXPLORE_MODEL |
gpt-6-luna |
Codex model for explore, read_slice, and web_lookup when neither the call nor the tool-specific _MODEL sets one. run_filtered defaults to the fast_delegate cascade, not this. |
POLYAGENT_EXPLORE_EFFORT |
medium |
Explicit Codex effort for explore, read_slice, and web_lookup when the call omits effort; ignored for non-Codex engines and does not set the run_filtered cascade. |
POLYAGENT_<TOOL>_ENGINE |
cascade | Per-tool engine for the four auxiliary tools — <TOOL> is EXPLORE, READ_SLICE, RUN_FILTERED, or WEB_LOOKUP. The call's own engine parameter beats it. Unset, all four use the fast_delegate cascade (codex first), skipping engines that are missing, unhealthy or out of quota, and engines that lack what the tool needs. Refused when the engine lacks what the tool needs: read-only (explore/read_slice/web_lookup, which outside codex comes from the sandbox) or web search (web_lookup, codex or claude). |
POLYAGENT_<TOOL>_MODEL |
(see above) | Per-tool model, same four names. The call's model beats it. With a non-codex engine and no model set anywhere, the engine's own default model is used. |
POLYAGENT_AGENT_PATHS |
(off) | Additional :-separated roots for named agent personas, searched before project/home .claude/agents and ~/.claude/plugins. |
POLYAGENT_SANDBOX |
bwrap |
Isolates every engine in a bubblewrap sandbox with an empty $HOME, preventing global config, MCP servers, hooks, and skills from loading. Only auth, required engine state, and toolchains are bound in. Set off/0 to disable explicitly — with the sandbox off, the read-only tools (explore, read_slice, web_lookup) accept only the codex engine. A missing bwrap is a startup error, never a silent downgrade. |
POLYAGENT_FORCE |
(off) | If 1/true, force-enable non-interactive approval for Cursor, Claude, and OpenCode runs. |
POLYAGENT_TIMEOUT_MS |
1800000 (30 min) |
Per-call safety-net timeout (not a work budget). Execution tools (delegate/fast_delegate) also get a prompt note so the worker returns partial results before being killed. |
POLYAGENT_LOG |
(off) | Path to a JSONL file; when set, calls log {tool, outChars} for bridge_stats. Internal Jev attempts also write tool:"decide", engine:"jev", zero returned chars, and a decision record with choice, confidence, acceptance, fallback, latency, cost, and requested level for shadow. |
POLYAGENT_JEV_FANOUT |
(off) | Set to 1/true/on to ask Jev whether 2+ successful fan_out consensus outputs substantially agree. At or above the threshold, return the first worker and session handles without the Codex arbiter. On low confidence, invalid answer, missing key, Jev error, or a 5-second timeout, use the arbiter. Pay-per-token OpenRouter call. |
POLYAGENT_JEV_FANOUT_THRESHOLD |
0.85 |
Minimum Jev agreement probability for skipping the arbiter; invalid values use 0.85. |
POLYAGENT_CLAUDE_MD |
(on) | Set to off/0/false/no to stop the server from refreshing an installed CLAUDE.md block on start. |
POLYAGENT_CLAUDE_MD_PATH |
~/.claude/CLAUDE.md |
File that holds the managed block. |
POLYAGENT_JEV_SHADOW |
(off) | Set to 1/true/on to ask Jev for a delegate level and a prompt-clarity check (target named, done defined, single task) in one parallel call with the worker. The suggestion never changes the requested level; a pending Jev call adds at most 5 seconds after the worker ends. Pay-per-token OpenRouter call. |
POLYAGENT_HOOK_MODE |
redirect |
Hook behavior: off (no-op), nudge (non-blocking additionalContext only), or redirect (deny once + name bridge tool for WebSearch/WebFetch and whole-file large Read; fail-open on retry). Grep/Glob/Bash/Edit/Write stay nudge-only. |
POLYAGENT_HOOK_MIN_LINES |
300 |
Line threshold above which the optional hook (below) redirects/nudges whole-file Read toward read_slice. |
Both Jev flags are off by default and pass truncated worker output or task text through the credential scrubber before sending it to OpenRouter.
They use the same OPENROUTER_API_KEY or OpenCode auth-file key as decide.
Every CURSOR_BRIDGE_* variable was renamed to POLYAGENT_* (same suffix), and CURSOR_BIN
became POLYAGENT_CURSOR_BIN. This is a clean cut: the old names are no longer read at all —
setting one has zero effect (no fallback, no warning). Update your host config (mcp.json /
settings.json "env" blocks) and any shell profile before upgrading.
| Old (removed) | New |
|---|---|
CURSOR_BIN |
POLYAGENT_CURSOR_BIN |
CURSOR_BRIDGE_AGENT_PATHS |
POLYAGENT_AGENT_PATHS |
CURSOR_BRIDGE_GROK_BIN |
POLYAGENT_GROK_BIN |
CURSOR_BRIDGE_CODEX_BIN |
POLYAGENT_CODEX_BIN |
CURSOR_BRIDGE_CLAUDE_BIN |
POLYAGENT_CLAUDE_BIN |
CURSOR_BRIDGE_MODEL |
POLYAGENT_MODEL |
CURSOR_BRIDGE_EXPLORE_MODEL |
POLYAGENT_EXPLORE_MODEL |
CURSOR_BRIDGE_IMAGE_MODEL |
POLYAGENT_IMAGE_MODEL |
CURSOR_BRIDGE_FORCE |
POLYAGENT_FORCE |
CURSOR_BRIDGE_ENABLE_CURSOR |
POLYAGENT_ENABLE_CURSOR |
CURSOR_BRIDGE_TIMEOUT_MS |
POLYAGENT_TIMEOUT_MS |
CURSOR_BRIDGE_DEBUG |
POLYAGENT_DEBUG |
CURSOR_BRIDGE_SANDBOX |
POLYAGENT_SANDBOX |
CURSOR_BRIDGE_SANDBOX_EXTRA |
POLYAGENT_SANDBOX_EXTRA |
CURSOR_BRIDGE_LOG |
POLYAGENT_LOG |
CURSOR_BRIDGE_HOOK_MODE |
POLYAGENT_HOOK_MODE |
CURSOR_BRIDGE_HOOK_MIN_LINES |
POLYAGENT_HOOK_MIN_LINES |
"cursor" survives only where it names the actual Cursor engine (POLYAGENT_CURSOR_BIN,
POLYAGENT_ENABLE_CURSOR).
Security:
delegate,fast_delegate, andrun_filteredhave full access and auto-approve their work.explore,read_slice, andweb_lookupuse Codex's read-only sandbox. Named agents are resolved on the host, reject path traversal, and are injected without mounting agent directories.
Managed CLAUDE.md block. Run npm run install-claude-md -- --dry-run to preview, then
npm run install-claude-md (or npx -p polyagent-mcp polyagent-mcp-install-claude-md). It appends a
routing block to ~/.claude/CLAUDE.md between <!-- polyagent-mcp:begin … --> and
<!-- polyagent-mcp:end -->, replaces an old hand-written <polyagent_preference> section, and saves
CLAUDE.md.polyagent.bak. From then on the server refreshes the block on each start (only when the
markers exist). Set POLYAGENT_CLAUDE_MD=off to stop that, or delete both markers.
Registering the tools is not enough. Two structural forces push the agent back to
native tools: (1) the host rule "prefer the dedicated file/search tools", and (2) MCP
tools used to be deferred — the agent had to run a tool-search to load their schemas,
so always-loaded Read/Grep/WebSearch won by default. The server now publishes
startup instructions (routing boundary) and marks the five core tools, fast_delegate,
fan_out, and rate with
_meta: { "anthropic/alwaysLoad": true } (Claude Code ≥2.1.121) so their schemas load
eagerly. Both fast_delegate and fan_out joined them after they were deferred and never got
called — the same adoption bug. The remaining secondary tools (generate_image, follow_up,
bridge_stats, decide) stay deferred. Four fixes, strongest first:
1. Call-time hook (recommended). A PreToolUse hook that steers the agent toward
the bridge at the moment it reaches for a native tool — text in a config file loses under
pressure, a call-time reminder does not. This repo ships one at
hooks/prefer-polyagent.mjs: it runs on node
(already required) and only fires where it pays. Default mode is redirect
(POLYAGENT_HOOK_MODE=redirect): for the two safe-to-block cases it returns
permissionDecision: "deny" once and names the bridge tool; other cases stay non-blocking
nudges. Wire it into your host's settings (Claude Code settings.json):
Breaking change (US-007): The hook file was renamed from
hooks/prefer-cursor-bridge.mjstohooks/prefer-polyagent.mjs. Update any hostsettings.jsonentry that points to the old path.
{
"hooks": {
"UserPromptSubmit": [
{
"hooks": [
{ "type": "command", "command": "node /abs/path/to/polyagent-mcp/hooks/prefer-polyagent.mjs", "timeout": 5 }
]
}
],
"PreToolUse": [
{
"matcher": "Read|Grep|Glob|WebSearch|WebFetch|Bash|Edit|Write|MultiEdit",
"hooks": [
{ "type": "command", "command": "node /abs/path/to/polyagent-mcp/hooks/prefer-polyagent.mjs", "timeout": 5 }
]
}
]
}
}UserPromptSubmit runs on every submitted prompt. That event has no matcher field; the hook's
promptRouteContext() detects routing cues and stays silent when none match. This path is
intentionally not deduplicated: each prompt is judged on its own.
For PreToolUse, each nudge/redirect fires at most once per session (deduplicated in a tmp
file keyed by session_id), because a repeated fire is worse than none: the agent learns
to ignore it and every fire costs tokens. Dedup keys are saved before emitting so
redirect is one-shot and fail-open (a second identical call is allowed through).
Readwhole-file (no offset/limit) overPOLYAGENT_HOOK_MIN_LINESlines → redirect (default) or nudge towardread_slice(once per file). Partial reads are left alone.WebSearch/WebFetch→ redirect (default) or nudge towardweb_lookup(once).Grep/Glob→ emits the one-time preload reminder to run theToolSearchfor any still-deferred bridge tools (nudge only — never redirected). The dedup collapses them to a single fire.Bashwhose command writes an artifact (git commit/push,git worktree add,gh pr create,gh issue create,bkt pr create) → suggests offloading that grunt-work todelegate(once, nudge only). Read-only Bash (status/diff/log/checkout) is left alone — the orchestrator needs that state, and a mechanical filter (e.g. rtk) already trims the noise.Edit/Write/MultiEdit→ once per session, reminds that a self-contained task (feature, bugfix, mechanical multi-file change, build fix) can go whole todelegate(prompt, level)— the selected worker edits with full access — instead of the orchestrator implementing it on expensive tokens. It never blocks the edit (nudge only); the once-per-session dedup means the orchestrator still edits inline freely (the nudge repositions execution, it doesn't police every edit).- The first qualifying fire of the session (whichever tool triggers it) also carries that preload reminder, so secondary schemas get loaded even in a Read-only or web-only session.
- Redirect deny reasons end with a fail-open suffix: if the bridge tool isn't loaded yet, run ToolSearch first; if the native tool is genuinely needed, call it again and it will be allowed (critical under headless
-pso the agent never hard-stalls).
Set POLYAGENT_HOOK_MODE=nudge for the old non-blocking behavior, or off to disable.
To reset the dedup and see the fires again, start a new session (or delete
polyagent-nudged-<session_id>.json from your OS temp dir — os.tmpdir(),
e.g. /tmp on Linux, not necessarily $TMPDIR).
The PreToolUse preload above only fires when the agent uses the Grep/Read tool. But
under pressure agents often reach for Bash grep instead, which matches no PreToolUse
matcher — so the preload reminder never arrives. Wire the same hook for SessionStart to
close that hole: the preload reminder then lands in context before the first tool decision,
regardless of how the agent searches.
{
"hooks": {
"SessionStart": [
{ "hooks": [{ "type": "command", "command": "node /abs/path/to/polyagent-mcp/hooks/prefer-polyagent.mjs", "timeout": 5 }] }
]
}
}On SessionStart the hook emits the ToolSearch preload as additionalContext and pre-marks
preload as seen in the session's dedup file, so the PreToolUse piggyback never repeats it.
The nudges above only steer the main loop. Spawned subagents never see them,
so wire the same hook for SubagentStart as well:
{
"hooks": {
"SubagentStart": [
{
"hooks": [
{ "type": "command", "command": "node /abs/path/to/polyagent-mcp/hooks/prefer-polyagent.mjs" }
]
}
]
}
}On SubagentStart the hook injects a compact polyagent preference into every
spawned subagent via additionalContext (subagentStartContext(agent_type)).
When agent_type is Explore it appends an extra line: that Explore run was spawned on the
orchestrator's expensive model (Explore inherits the session model, capped at Opus), so it should
route all reading through explore/read_slice (which run on GPT-6 Luna) and
keep the expensive shell to orchestration only.
Coexisting with context-mode. The bridge and context-mode use separate channels (context-mode may still do its own thing; this hook only emits
additionalContext), so they coexist cleanly — noupdatedInputrace, no delay, no import of context-mode's routing.
2. Load deferred tools when needed. The five core tools, fast_delegate, fan_out, and
rate are alwaysLoad on Claude Code ≥2.1.121. Older hosts may still defer their schemas. The
secondary tools (generate_image, follow_up, bridge_stats, decide) remain deferred; load only
the ones a task needs. Add to your CLAUDE.md/AGENTS.md:
If a required core schema is missing on an older host, run tool-search for
`delegate, fast_delegate, explore, read_slice, run_filtered, web_lookup, fan_out, rate`.
For secondary tools, load `generate_image`, `follow_up`, `bridge_stats`, or `decide` only when needed.
3. Reconcile the conflict in CLAUDE.md. State the precedence explicitly:
The host rule "prefer dedicated file/search tools" applies to the EDIT path (Edit needs the
file content → native Read). For PURE reading/locating/web (no edit), polyagent takes
precedence over native Read/Grep/Glob/WebSearch/WebFetch. Read a large file whole with native
Read ONLY when you are about to edit it.
4. Delegate execution, not just exploration. The bridge is not only for reading — delegate
runs implementation work with full read/edit/shell access, so the orchestrator shouldn't burn its
own tokens on self-contained tasks. State this in CLAUDE.md so the agent routes doing, not just
finding, to the cheap worker:
You are the ORCHESTRATOR. delegate(prompt, level) is the DEFAULT for BOTH execution AND judgment.
`level` picks a distinct tier: 1=GPT-6 Luna max (codex), 2=GPT-6 Sol high (codex), 3=GPT-6 Sol
max (codex), 4=Claude Opus 5.5 high (claude), 5=Claude Opus 5.5 max (claude). Levels 4 and 5 are expensive
(5 costs ~3.3× level 4) — reserve them for what cheaper levels cannot do. Both use the Claude Code
host subscription: level 4 is 44% cheaper than before but now spends host quota. Codex holds levels
1–3; Claude holds 4–5, so exhausted Claude quota takes down both levels and that host together. The worker has full read/edit/shell access
in cwd when you delegate. The constant win is context economy: the worker's raw output never enters
your context. Delegate it, then review the result; edit inline only for a quick one-off you're
already positioned for. Use fast_delegate(prompt) when the work is self-contained and you just want
the fastest healthy worker. Pass agent:"name" or agent:{prompt:"..."} when the worker
needs a specialized persona.
The AA cost ladder (quota proxies, not per-call bills) is $0.07 → $0.37 (~5×) → $1.06 (~3×) → $1.82 (~1.7×) → $5.98 (~3.3×). See the level 4 decision.
npm test # vitest — unit tests for model resolution / arg building
npm run dev # run from source via tsxMIT