Open source (MIT) · v0.10.1 · public beta

Six pinned specialists. One lead you control.

Otter Swarm turns Devin CLI and Devin Desktop into a team. A planner, researcher, implementer, reviewer, tester and interrogator — each pinned to the exact model you choose, each returning evidence, every run leaving a durable record.

Built for Devin CLI and Devin Desktop
devin — local session
0 / 417 unit tests green (npm test)
0 e2e scenarios · 0 checks passing
0 pinned roles
$0 implementer cost on the free SWE-2 pin (bench, 2026-09-13)

Measured on the Standard preset; free-tier pricing is promo-dependent.

01 / How it works

Your session. Six workers on rails.

The lead stays yours

/model picks the lead. It talks to you, asks the questions, and runs all web tooling.

Workers are pinned, persistently

~/.config/devin/agents/<role>.md holds an exact uid plus a fallback chain, verified against your account's model list. Pins survive sessions — and are checked at dispatch: a profile whose model: disagrees with the recorded pin is refused before the worker spawns, not silently substituted.

Work is bounded by contracts

Owns: paths, Touched / Ran / Next footers, P0–P3 findings with evidence, PASS / FAIL / INCONCLUSIVE gates. A missing lane never approves.

Every run is auditable

Run manifests in .devin/otter/runs/, attest blocks on the plan, and /otter:stats record what actually happened.

lead your /model planner researcher implementer reviewer tester interrogator

Briefs go out. Evidence comes back: Touched / Ran / Next.

02 / Roles

Six specialists, each with a lane

planner

Read-only planning; returns questions, writes nothing.

read-only

researcher

Repo and source-pack analysis; read-only map.

read-only

implementer

One exact Owns:-bounded write slice, test-first.

write, bounded

reviewer

Clean-context P0–P3 critique of a diff or plan.

read-only

tester

Runs the relevant tests; high-assurance acceptance matrices.

exec

interrogator

Opt-in adversarial verification of a change or plan.

read-only
Five cartoon otters as the swarm: one pointing, one with a blueprint, one in a hard hat, one thinking, one with a magnifier

The swarm: planner, researcher, implementer, reviewer, tester, interrogator.

03 / Pins

Pin every worker to the model you want. On disk. Forever.

A pin is an exact uid written to ~/.config/devin/agents/<role>.md and verified against devin models list. No per-session overrides, no invented ids.

  • Pins are exact uids in ~/.config/devin/agents/<role>.md.
  • /otter:setup always asks: Standard / Budget / Max / pick each role.
  • Every uid is verified against devin models list.
  • Each role carries a fallback chain.
  • The resolver is fail-closed — the error names the role it could not resolve.
  • It never invents a model id.
  • /otter:models reviewer grok-4-6-high changes exactly one file.
  • Preset applies stamp only the model: line — your hand-edited prompt body survives.
  • /otter:models off / on park and restore the whole panel.
day-2 model management
/otter:models                 show current pins
/otter:models standard|budget|max
/otter:models reviewer grok-4-6-high
/otter:models repair
/otter:models off · on

Standard's reviewer shares a family with the implementer. For a cross-family opinion pin a different reviewer or run /otter:review dual. At standard assurance the lead itself reviews; review=worker brings the pinned reviewer back.

04 / Pipeline

/otter:orchestrate — evidence in, verdict out

The lead sizes the work, asks, then runs bounded lanes. Every lane leaves a receipt — and the plan gate applies at every assurance level: you approve the plan before any writer starts.

authsecretspaymentsmigrationspublic APIsconcurrencyCI/deploy>3 writer slices any one forces high
1

Size & clarify

Classify the task, ask 1–3 questions, require exact Owns: paths before anything spawns. You see the whole plan and approve, revise, or abort before any writer starts.

lead
2

Research

One researcher by default; fan-out of 2–3 only for independent questions. The lead builds source packs.

researcher
3

Implement

One writer per tree, Owns:-bounded, test-first. Independent slices run in parallel worktrees — the plan itself says which: each slice declares Covers: (the criteria it satisfies) and Depends on:. Edits outside Owns are flagged and reviewed by the lead, not bounced to you.

implementer
4

Test

The lead re-runs the green test command; a tester matrix runs on high or when evidence is missing.

tester
5

Review

The lead reviews the diff under the reviewer contract — P0–P3 findings, each backed by a probe, validated by script. On high assurance or review=worker a clean-context reviewer runs instead. An open P0 or P1 is a FAIL.

lead · reviewer on high
6

Repair

Only request_changes, a failed test lane or an open P0/P1 starts a repair round; an approve with P2/P3 findings ends the run and lists them as follow-ups. One sequential implementer repairs; the loop never runs blind.

implementer
7

Prove

Git diff, exit codes, artifacts. Worker summaries are not proof. The acceptance criteria the evidence covers get ticked in the plan; the reply shows AC → evidence.

lead
8

Attest

Timestamp, panel-observed profile names and requested pins written to the run and the plan.

lead

Parallel writers by default — disjoint turf, one tree each

Slices with disjoint Owns: fan out automatically into linked worktrees — up to four writers at once, never two on one tree. Binary patches are applied back to main sequentially and hash-checked, then tested and reviewed once on the integrated tree. worktrees=1 forces sequential.

HEAD worktree 1 · writer 1 Owns: src/auth/* worktree 2 · writer 2 Owns: src/api/* worktree N · writer N ≤ 4 Owns: disjoint paths main tree binary patches applied sequentially · hash-checked · first conflict stops

An assurance instrument, not a speed hack

On a 5-module task the pipeline ran 6× the wall time and 2.9× the cost of a stock lead — down from 20× and 6.6× three days earlier, after fixing our own tooling friction, moving state writes to one checkpoint per phase, and letting the lead review at standard assurance. Single-worker skills like /otter:implement passed every gate at $0 on the free SWE-2 pin.

bench (2026-09-13)passcost
implement · 1 file5/5$0 vs $0.06
implement · 5-module5/5$0 vs $0.19
orchestrate · 5-module5/52.9× ($0.59)
review · planted off-by-one5/52.4×
review · 3 planted bugs5/5 caught (stock too)1.3×
review dual · 3 planted bugs5/5 caught (stock too)4.5×
research · repo question5/51.2×

Honest note: on a standalone diff a strong lead found the same three bugs in 30 s. The reviewer's value is independence inside orchestrate and the validated contract, not raw detection; dual is for high assurance.

See the full run, step by step →Diagram, state writes, tests, benefits and the caveats — one page.

05 / Research

Research that cites, or says why it can't

/otter:research

One researcher. Classifies the question repo / web / mixed. For web, the lead fetches up to 8 sources into a source pack — subagents have no web tools.

/otter:research … wide

Up to 3 parallel researchers on cleanly decomposable questions. Disagreements are kept, not averaged. output=<path> writes cited markdown.

/otter:deep-research experimental

Plan → Research (≤4 researchers, lead-fed source packs) → Verify (2 reviewer shards, independent corroboration from a second source) → Report. Cited markdown; status Verified / Source-checked / Partial. Budget N+2 calls nominal, 2N+4 worst.

lead source packs researcher researcher researcher ≤4 verifier 1 verifier 2 report [S1]…[Sn] Verified Source-checked Partial

06 / Reports

Every run leaves a record

Run manifests

.devin/otter/runs/<run-id>.json — workers, requested pins, phases, lanes, budgets, outcome.

Attest blocks

Appended to the plan file: timestamp, panel-observed profile names, run-time pins.

Usage report

/otter:stats report week writes reports/otter-usage-<scope>.html — a self-contained dashboard: the data travels inside the file, so period, project, agent and model filters, the timeline and per-session drill-down all run in the browser, zero network requests. Cost from a live price snapshot, --redact for sharing, stats serve runs it live.

Gate receipts

/otter:gate aggregates lanes into PASS / FAIL / INCONCLUSIVE. A missing lane never approves.

Deep-research reports

Cited markdown with [S1]…[Sn] markers and honest coverage notes.

E2E + efficiency reports

node scripts/e2e.mjs --report and --bench keep the whole suite honest.

otter-telemetry MCP server

The same data as in-session tools, answerable from any session — even ones that never invoke an otter skill.

"what did I burn today?" "attribute this week by role" "generate this month's usage report"
Otter Swarm — usage dashboard · week illustrative
9sessions
1.8kturns
412ktokens
$1.23est. cost
lead
implementer
researcher
housekeeping
lead the session's own agent otter worker a pinned role that actually ran other subagent sidekick, explore, custom profiles housekeeping summarizer / compactor

Attribution is structural: each turn is filed under the agent whose system prompt roots its subtree, so every turn lands in exactly one bucket — nothing sits in "unattributed".

07 / Command reference

Which command do I run?

Everyday chat is normal Devin on your picker model. Otter Swarm only runs when you call an /otter: command.

Pick by what you want

You wantRun
A second opinion on a diff/otter:review (add dual for reviewer + interrogator)
A feature built with verification/otter:orchestrate <task>
To plan before committing to a buildDevin /plan, leave plan mode, then /otter:orchestrate
Research on the repo/otter:research <question> (wide for parallel)
Research on the web/otter:deep-research <query> (experimental)
A bug found and fixed properly/otter:diagnose <symptom>
A plan stress-tested before building/otter:hyperplan <plan> or /otter:plan review
Proof a change meets its acceptance criteria/otter:gate plan=<path>
Work that survives session interruptions/otter:goal set <objective>
To know what a session cost/otter:stats — or just ask "what did I burn today?"
A weekly usage report/otter:stats report week

Set up

Once per machine. Pins live on disk and survive sessions.

/otter:setup lead only

What it doesWrites the six worker profiles and the pin sidecar. Always asks how you want worker models: Standard, Budget, Max, or pick each role.

When to useFirst run. After an update, run /otter:doctor instead.

/otter:setup

/otter:models … lead only

What it doesShow or change worker pins: a whole preset, one role at a time, repair to re-verify against your account, off / on to park or restore all profiles.

When to useYou want a different model for one role, or a worker failed to spawn.

/otter:models reviewer grok-4-6-high

/otter:doctor [verbose|json|live] lead only

What it doesStatic checks of the installation, profiles, pins, state, and known platform limits. A paid dispatch smoke test only runs with your explicit confirmation.

When to useSomething stopped working, especially after devin plugins update.

/otter:doctor verbose

/otter:beta-report lead only

What it doesPacks a redacted diagnostics bundle — doctor output, the usage cube, run manifests, price-snapshot status — into .devin/otter/beta/ plus a .tgz. Nothing is uploaded; you attach the archive to an issue yourself.

When to useReporting a bug during the beta — the bundle shows what ran without sharing your code or prompts.

/otter:beta-report

Build

One writer per tree. Every slice declares the files it may touch; disjoint slices run in parallel.

/otter:plan <objective> workers: planner

What it doesThe read-only planner drafts a canonical plan and returns owner questions. You answer, the lead validates and writes the approved plan.

When to useYou want scope settled before anything is built.

/otter:plan add rate limiting to the API gateway

/otter:orchestrate <task> [assurance=light|standard|high] [review=lead|worker] [worktrees=1|N] [plan=<path>] workers: implementer · researcher, tester, reviewer by assurance

What it doesSizes and clarifies the work, then runs research when needed, implementers on Owns:-bounded slices (disjoint slices fan out to parallel worktrees), the lead's own test re-run and review at standard assurance (tester matrix and reviewer worker on high), with durable run state written once per phase.

When to useA feature built with verification.

/otter:orchestrate add rate limiting to the API gateway assurance=high

/otter:implement … workers: implementer

What it doesOne implementer on one exact Owns:-bounded slice, test-first. A slice that needs an unowned file edits it minimally and flags it; the lead reviews the whole diff.

When to useYou know exactly which files change and want a cheap pinned writer.

/otter:implement Owns: src/limiter.mjs tests/limiter.test.mjs — add a token bucket

/otter:diagnose <symptom> workers: implementer, tester

What it doesRed loop first: falsifiable hypotheses, one-variable probes, a regression that goes red then green, the fix, and every debug marker removed.

When to useA bug you want found and fixed properly.

/otter:diagnose tests fail after the migration refactor

Verify

Clean-context critics. Findings come with evidence, never fixes you did not approve.

/otter:review [dual] workers: reviewer (+ interrogator)

What it doesA reviewer that has not seen your reasoning returns P0–P3 findings with evidence and a verdict. dual adds the interrogator in parallel and merges a triaged shortlist.

When to useA second opinion on a diff — especially one the lead itself wrote. Measured: on a foreign diff a strong lead finds the same bugs; the worker buys a clean context and a validated contract.

/otter:review dual

/otter:test workers: tester

What it doesRuns the relevant tests, or a high-assurance acceptance matrix with criterion, command, expected and observed per row.

When to useYou need test evidence, not a summary.

/otter:test

/otter:interrogate workers: interrogator

What it doesStandalone adversarial verification of a change or plan: assume it is wrong and try to prove it. Opt-in; not part of the default pipeline.

When to useHigh-stakes changes.

/otter:interrogate

/otter:hyperplan <plan|diff> workers: interrogator ×2–4

What it doesTwo to four orthogonal interrogator critics on one plan or diff. The synthesis dedupes, ranks severity, and preserves contradictions.

When to useA plan stress-tested before building.

/otter:hyperplan plans/rate-limit.md

/otter:gate plan=<path> workers: tester, reviewer

What it doesAggregates acceptance, test and review lanes into exactly one verdict: PASS, FAIL or INCONCLUSIVE. A missing lane never approves.

When to useProof a change meets its acceptance criteria.

/otter:gate plan=plans/rate-limit.md

Research

Read-only. Web sources are fetched by the lead; workers analyze.

/otter:research <question> [wide] [output=<path>] workers: researcher (×3 with wide)

What it doesA read-only researcher answers repo questions directly; for web questions the lead fetches up to 8 sources into a source pack first. wide fans out up to 3 researchers and keeps disagreements.

When to useResearch on the repo.

/otter:research where is auth token refresh handled?

/otter:deep-research <query> [breadth=N] [output=<path>] [verify=quick] workers: researcher ×≤4, reviewer ×2

What it doesPlan, then up to 4 researchers, then 2 verifier shards with independent corroboration, then a cited markdown report marked Verified, Source-checked or Partial. Experimental.

When to useResearch on the web, with citations.

/otter:deep-research compare rate-limiter libraries for Node

Track

Every run is durable and local.

/otter:status [json] lead only

What it doesBounded local run and goal status: phase, workers, requested pins, budgets, blockers, next action.

When to useYou want to know where a run stands.

/otter:status

/otter:resume <run-id> lead only

What it doesResumes the first safe incomplete phase of an interrupted run. Completed or unknown writes are never repeated automatically.

When to useA session was interrupted mid-run.

/otter:resume run-20260912-4f1c9a2b7e0d

/otter:goal … lead only

What it doesOne persistent project objective: approved plan, linked runs, call and round budgets, pause and resume, evidence-gated completion.

When to useWork that has to survive session interruptions.

/otter:goal set ship v2 of the parser

/otter:stats [session <id>] [week|month|since <Nd>] [report] lead only

What it doesWhat a session or period burned, per agent and per model — every turn attributed by conversation tree. report writes a self-contained dashboard under reports/; serve runs it live.

When to useYou want to know what a session cost.

/otter:stats report week

08 / Safety

Guardrails, not vibes

The lead walks in. Then the swarm.

One writer per tree

Independent slices run in parallel, one implementer per linked worktree, hash-checked patches, one verification on main. Never two writers on one tree.

Owns bounds

A slice needing an unowned file edits it and flags it under Outside Owns; parallel worktree mode still stops instead.

Command guard

An optional PreToolUse hook offered at setup blocks rm aimed at ~ or /, force-push, keychain access and pipe-to-shell — even inside chained commands. Fails closed while ~/.config/devin/otter-guard exists.

Local, read-only telemetry

The sessions database is opened read-only. Nothing is uploaded.

Fail-closed pins

No id in the chain resolves on your account → nothing is written, and the error names the role. Drift later? The dispatch check refuses the spawn and names /otter:models repair.

09 / Get ready

Ready your rig

Requirements

  • Devin CLI or Devin Desktop local agent
  • Node 18+ for the helper scripts; telemetry needs Node ≥ 22.12
  • macOS or Linux — Windows is untested in 0.10.x
  • Devin plugins enabled on your account (devin plugins --help)
  • A plan whose preset chains resolve on your account
path A — profiles only
mkdir -p ~/.config/devin/agents
cp templates/agents/*.md ~/.config/devin/agents/
path B — full plugin
devin plugins install chris-wozniczek/otter-swarm # public — plugin beta access required
/otter:setup                                      # pick a preset, opt into the guard + update check

devin plugins update otter                        # later: pull a release, then start a new session and run /otter:doctor

Updating

devin plugins update otter pulls a release, then start a new session — skill bodies load at session start — and run /otter:doctor. Pins, profile bodies and run state live outside the plugin root and are never touched. The opt-in session-start check tells you once per new version: one GET of plugin.json, nothing sent — touch ~/.config/devin/otter-no-update-check to stop.

Source

MIT-licensed, public on GitHub — issues, the beta-report template, ATTRIBUTIONS.md, and the full test log live in the repository.

GitHub