Ask
youYou type /otter:orchestrate <task>. Plain chat never triggers the pipeline.
Deep dive · v0.10.1
What the lead does, what the workers do, what lands on disk, and how we know. Read it top to bottom or jump to the caveats — they are the honest part.
01 / Diagram
Cyan packets are briefs going out; amber packets are evidence coming back. Every hand-off crosses the disk lane through a checkpoint — the run manifest is written before anything is shown to you.
02 / Steps
You type /otter:orchestrate <task>. Plain chat never triggers the pipeline.
Classifies the request (review-only and test-only route to their own skills), picks an assurance level from observable signals and states it in one line. If goal, constraints, non-goals, acceptance criteria or scope are unclear it asks 1–3 questions and spawns nothing.
evidencethe assurance line, e.g. "assurance: light — explicit Owns, npm test present, no risk signals"
Writes or validates a canonical plan: goal, verified context, constraints, non-goals, stable AC IDs, approach, reuse, Scope In/Out, ordered slices with exact Owns paths plus per-slice Covers and Depends on, risks, verification. "The auth package" is not a path. Then the plan gate: you see the whole plan and approve, revise, or abort — run-state init refuses a missing or stale approval, so no writer starts without it.
evidenceplans/<slug>.md passing plan-contract validate
One run-state init call validates the plan against the contract, compares the implementer profile's model: with the recorded pin, then in a single locked write creates the manifest, closes planning, opens implementation and assigns the worker attempt — printing the run id, plugin root and attempt id the next commands need. A bad plan or drifted pin refuses before anything is written; a bad pin never reaches a spawn.
evidence.devin/otter/runs/<run-id>.json with phase implementation and attempt-implementer-1
Standard: only when the lead still has an open question after reading the owned files. Fan-out to 2–3 lanes only for genuinely independent questions, after showing the cost. Subagents have no web tools — the lead fetches sources into packs. No writer starts until every lane settles.
evidencemerged findings with paths/URLs, contradictions kept, union of "Not searched"
One writer per tree, Owns-bounded, test-first — slices with disjoint Owns fan out to parallel linked worktrees by default. On spawn the worker runs a one-command preflight.mjs self-check — its own pin against the recorded pin — as defence-in-depth under the dispatch-time gate. Edits outside Owns are made minimally and flagged under Outside Owns; the lead reviews them in the diff.
evidenceimpl.diff; the worker's Touched / Ran / Next footer
The lead re-runs the implementer's own green test command once and records the exit code as the test lane — it does not dispatch a tester to repeat it. A tester matrix runs on high assurance or when there is no test evidence.
evidencenpm-test.txt with exit=0; lane test: pass
The lead applies the reviewer contract to itself: reads the whole diff and the code around it, runs one probe per suspicion, writes P0–P3 findings with location, trigger, impact and evidence, validated by script. Open P0/P1 or request_changes is FAIL. Malformed after one retry is INCONCLUSIVE — never approval.
evidencereview.md passing evidence-contracts review; lane review with "lead review under reviewer contract"
Only a request_changes verdict, a failed test lane or an open P0/P1 starts a repair round. An approve with P2/P3 findings ends the run and lists them as follow-ups. One sequential implementer repairs; only affected lanes re-run; the loop never runs blind.
evidenceattempt-implementer-2; re-review
The lead reads the actual git diff, exit codes and artifacts — worker summaries are not proof. It ticks the acceptance criteria the evidence covers (plan-contract.mjs prove; an unticked AC at completion is FAIL unless you waive it), completes the run, appends an attest block (timestamp, profile names, requested pins, profile digests) to the plan, and replies in short sentences: assurance level, what was skipped, changed paths, verification, AC → evidence, run id, next command.
evidencemanifest status completed · outcome pass|fail|inconclusive
node $ROOT/scripts/run-state.mjs init --kind orchestrate --assurance standard --plan plans/plan.md --evidence plans/plan.md --next implementation --start implementer
node $ROOT/scripts/run-state.mjs checkpoint --phase implementation --finish attempt-implementer-1 --evidence impl.diff,npm-test.txt --lane test --outcome pass --next review
node $ROOT/scripts/run-state.mjs checkpoint --phase review --evidence review.md --outcome pass --complete
# high: --start reviewer · light: --lane review --summary "waived: light assurance"
03 / State
.devin/otter/runs/<run-id>.json — kind, assurance, plan digest, phases with idempotency keys, workers (attempt id, role, requested pin, profile digest, status incl. blocked, owns_outside), lanes with evidence paths and summaries, budget, blockers, outcome.
.devin/otter/evidence/ — impl.diff, npm-test.txt, review.md, test.md — the files the lanes point at. A lane whose evidence file is missing cannot be recorded.
plans/<slug>.md — The deliverable you may commit. The attest block at the end records what ran and the no-model-telemetry limitation in plain words.
.devin/otter/ is added to .git/info/exclude on first write; plans/ stays visible on purpose. The first write also stamps .devin/otter/plugin-root, so later commands resolve the plugin from the workspace instead of re-asking the CLI. /otter:status and /otter:resume read the manifest; /otter:stats reads sessions.db.
04 / Assurance
| lane | light | standard | high |
|---|---|---|---|
| researcher | — | only with open questions | yes |
| implementer | yes (one per tree) | yes (one per tree) | yes (one per tree) |
| test | lead re-runs repo command | lead re-runs the green command | tester matrix vs every AC |
| review | waived, recorded | lead under reviewer contract (review=worker dispatches) | reviewer worker; interrogator on request |
| typical dispatches | 1 | 1 | 4–5 |
| forced by | explicit Owns, test command, ≤3 files, no risk signal | default | any risk signal, or assurance=high |
Skipping a lane never removes its evidence requirement — a light run records review: pass with summary "waived: light assurance" so the manifest states what was not done. A goal's assurance wins over the lead's guess and over your flag.
05 / Tested
Every claim above is backed by a headless e2e session: real devin -p runs in throwaway git workspaces, user config snapshotted and restored, graded against a deterministic contract — files, manifests, the sessions database, never a model's opinion. 2026-09-14/15, Standard preset, kimi-k3-high lead.
| scenario | what it proves |
|---|---|
session-orchestrate | plan-driven run wrote the manifest, phases ran in order, deliverable produced |
session-orchestrate-sizing-small | one-file change sized light; no researcher spawn; implementer ran; npm test exit 0 |
session-orchestrate-ambiguity | "make it better" asked questions and spawned nothing |
session-six-profiles | researcher, implementer, reviewer and tester each ran under their own otter:<role> name; assurance = high |
session-owns-outside | an unowned test file was created and flagged under Outside Owns; suite green |
session-worktree-owns-stop | a cross-Owns edit in worktree mode stopped the worker (blocked); the lead recovered sequentially on main |
session-research-fanout | three researcher lanes on a decomposable question, merged with contradictions and "Not searched" |
session-bad-pin-garbage | a corrupted model: line was refused at run open (init/checkpoint --start); zero implementer turns; target file byte-identical |
session-resume-interrupted | an interrupted run resumed at the first safe incomplete phase |
session-worktrees | two writers in linked trees, patches applied sequentially and hash-checked |
Full table and evidence pointers: E2E-REPORT.md and TEST.md in the repository. Two checks stay human-only: the guard hook's live block (needs a Desktop relaunch) and a background worker's permission denial (needs the interactive prompt).
06 / Benefits
The implementer never sees your conversation; the reviewer never sees the implementer's reasoning. Independence is structural, not requested.
Owns paths, one writer per tree, linked worktrees for disjoint slices with hash-checked patches. A worker can edit an unowned file when it must — and the lead sees it flagged in the diff.
Manifest, evidence files, attest block, usage dashboard. Who ran, on which requested pin, what was verified, what it cost.
Free or cheap pins for the roles that read and test, stronger pins where judgement matters. /otter:implement on the free SWE-2 pin passed every bench gate at $0.
A missing profile, a drifted pin, a missing evidence file or a malformed review stops the run and names the fix. Nothing silently downgrades.
07 / Caveats
Orchestrate is an assurance instrument for work large enough that a single context degrades. It is not a way to finish small tasks faster.
5-module feature, 5 trials, 2026-09-13: 6.0× the wall time and 2.9× the cost ($0.59) of a stock kimi-k3-high lead, ~21 lead turns, one dispatch at standard assurance. Two-slice runs last measured at 20–22× / 11–12× on 0.9.0. Single-worker skills are the sweet spot.
The implementer is told to stay inside its paths and to flag anything else; only worktree mode fails closed mechanically. The lead's diff read is what catches a misbehaving model.
Direct /otter:implement and the other role skills are dispatched by the platform with no script seam; a corrupted model: line there may be silently substituted. Profiles run a one-command preflight.mjs self-check on spawn — defence in depth, not enforcement. /otter:models repair fixes drift either way.
The manifest stores what was requested and the profile digest. The dashboard attributes turns by conversation tree from sessions.db and flags pin drift — that is telemetry, not a platform guarantee.
Every brief must be complete. Custom subagents get no web tools; the lead fetches sources into packs.
Install, update or edit → full relaunch (Cmd+Q). The slash dropdown empties past /otter: — type the full command.
Guard hook live block and background-worker permission denial are in the beta guide, not the harness.
That's by measurement (it caught the prototype-chain bug 5/5 vs the worker's 3/5) — but it means the second opinion at standard assurance is the same model that briefed the implementer. review=worker or assurance=high brings the clean-context worker back.