Deep dive · v0.10.1

One run of /otter:orchestrate, end to end

What the lead does, what the workers do, what lands on disk, and how we know. Read it top to bottom or jump to the caveats — they are the honest part.

01 / Diagram

Three lanes, one record

you lead · /model workers · pinned disk · .devin/otter repair · only fail / P0–P1 ask /orchestrate reply the receipt size & clarify assurance: light · standard · high plan test lead re-runs the green command review reviewer contract, probes, P0–P3 prove → attest researcher fan-out 2–3 implementer tester high / no evidence reviewer high · review=worker verdicts approve request_changes inconclusive plans/<slug>.md runs/<id>.json init pin ✓ impl.diff · npm-test.txt checkpoint run completed · attest block

Cyan packets are briefs going out; amber packets are evidence coming back. Every hand-off crosses the disk lane through a checkpoint — the run manifest is written before anything is shown to you.

02 / Steps

Ten steps, one writer per tree

01

Ask

you

You type /otter:orchestrate <task>. Plain chat never triggers the pipeline.

02

Size & clarify

lead

Classifies the request (review-only and test-only route to their own skills), picks an assurance level from observable signals and states it in one line. If goal, constraints, non-goals, acceptance criteria or scope are unclear it asks 1–3 questions and spawns nothing.

evidencethe assurance line, e.g. "assurance: light — explicit Owns, npm test present, no risk signals"

03

Plan

lead

Writes or validates a canonical plan: goal, verified context, constraints, non-goals, stable AC IDs, approach, reuse, Scope In/Out, ordered slices with exact Owns paths plus per-slice Covers and Depends on, risks, verification. "The auth package" is not a path. Then the plan gate: you see the whole plan and approve, revise, or abort — run-state init refuses a missing or stale approval, so no writer starts without it.

evidenceplans/<slug>.md passing plan-contract validate

04

Init the run

lead → disk

One run-state init call validates the plan against the contract, compares the implementer profile's model: with the recorded pin, then in a single locked write creates the manifest, closes planning, opens implementation and assigns the worker attempt — printing the run id, plugin root and attempt id the next commands need. A bad plan or drifted pin refuses before anything is written; a bad pin never reaches a spawn.

evidence.devin/otter/runs/<run-id>.json with phase implementation and attempt-implementer-1

05

Research

worker: researcher

Standard: only when the lead still has an open question after reading the owned files. Fan-out to 2–3 lanes only for genuinely independent questions, after showing the cost. Subagents have no web tools — the lead fetches sources into packs. No writer starts until every lane settles.

evidencemerged findings with paths/URLs, contradictions kept, union of "Not searched"

06

Implement

worker: implementer

One writer per tree, Owns-bounded, test-first — slices with disjoint Owns fan out to parallel linked worktrees by default. On spawn the worker runs a one-command preflight.mjs self-check — its own pin against the recorded pin — as defence-in-depth under the dispatch-time gate. Edits outside Owns are made minimally and flagged under Outside Owns; the lead reviews them in the diff.

evidenceimpl.diff; the worker's Touched / Ran / Next footer

07

Test

lead · tester on high

The lead re-runs the implementer's own green test command once and records the exit code as the test lane — it does not dispatch a tester to repeat it. A tester matrix runs on high assurance or when there is no test evidence.

evidencenpm-test.txt with exit=0; lane test: pass

08

Review

lead · reviewer on high

The lead applies the reviewer contract to itself: reads the whole diff and the code around it, runs one probe per suspicion, writes P0–P3 findings with location, trigger, impact and evidence, validated by script. Open P0/P1 or request_changes is FAIL. Malformed after one retry is INCONCLUSIVE — never approval.

evidencereview.md passing evidence-contracts review; lane review with "lead review under reviewer contract"

09

Repair (only when needed)

worker: implementer

Only a request_changes verdict, a failed test lane or an open P0/P1 starts a repair round. An approve with P2/P3 findings ends the run and lists them as follow-ups. One sequential implementer repairs; only affected lanes re-run; the loop never runs blind.

evidenceattempt-implementer-2; re-review

10

Prove, attest, reply

lead → disk → you

The lead reads the actual git diff, exit codes and artifacts — worker summaries are not proof. It ticks the acceptance criteria the evidence covers (plan-contract.mjs prove; an unticked AC at completion is FAIL unless you waive it), completes the run, appends an attest block (timestamp, profile names, requested pins, profile digests) to the plan, and replies in short sentences: assurance level, what was skipped, changed paths, verification, AC → evidence, run id, next command.

evidencemanifest status completed · outcome pass|fail|inconclusive

a standard run is one init and two checkpoints
node $ROOT/scripts/run-state.mjs init --kind orchestrate --assurance standard --plan plans/plan.md --evidence plans/plan.md --next implementation --start implementer
node $ROOT/scripts/run-state.mjs checkpoint --phase implementation --finish attempt-implementer-1 --evidence impl.diff,npm-test.txt --lane test --outcome pass --next review
node $ROOT/scripts/run-state.mjs checkpoint --phase review --evidence review.md --outcome pass --complete
# high: --start reviewer · light: --lane review --summary "waived: light assurance"

03 / State

What lands on disk

Run manifest

.devin/otter/runs/<run-id>.json — kind, assurance, plan digest, phases with idempotency keys, workers (attempt id, role, requested pin, profile digest, status incl. blocked, owns_outside), lanes with evidence paths and summaries, budget, blockers, outcome.

Evidence dir

.devin/otter/evidence/ — impl.diff, npm-test.txt, review.md, test.md — the files the lanes point at. A lane whose evidence file is missing cannot be recorded.

Plan + attest

plans/<slug>.md — The deliverable you may commit. The attest block at the end records what ran and the no-model-telemetry limitation in plain words.

Git hygiene

.devin/otter/ is added to .git/info/exclude on first write; plans/ stays visible on purpose. The first write also stamps .devin/otter/plugin-root, so later commands resolve the plugin from the workspace instead of re-asking the CLI. /otter:status and /otter:resume read the manifest; /otter:stats reads sessions.db.

04 / Assurance

Every lane, by level

lanelightstandardhigh
researcher—only with open questionsyes
implementeryes (one per tree)yes (one per tree)yes (one per tree)
testlead re-runs repo commandlead re-runs the green commandtester matrix vs every AC
reviewwaived, recordedlead under reviewer contract (review=worker dispatches)reviewer worker; interrogator on request
typical dispatches114–5
forced byexplicit Owns, test command, ≤3 files, no risk signaldefaultany risk signal, or assurance=high

Skipping a lane never removes its evidence requirement — a light run records review: pass with summary "waived: light assurance" so the manifest states what was not done. A goal's assurance wins over the lead's guess and over your flag.

05 / Tested

How we know it does this

Every claim above is backed by a headless e2e session: real devin -p runs in throwaway git workspaces, user config snapshotted and restored, graded against a deterministic contract — files, manifests, the sessions database, never a model's opinion. 2026-09-14/15, Standard preset, kimi-k3-high lead.

0 / 417 unit tests
0 e2e scenarios · 0 checks
0 live sessions in the pipeline suite
0 open findings
scenariowhat it proves
session-orchestrateplan-driven run wrote the manifest, phases ran in order, deliverable produced
session-orchestrate-sizing-smallone-file change sized light; no researcher spawn; implementer ran; npm test exit 0
session-orchestrate-ambiguity"make it better" asked questions and spawned nothing
session-six-profilesresearcher, implementer, reviewer and tester each ran under their own otter:<role> name; assurance = high
session-owns-outsidean unowned test file was created and flagged under Outside Owns; suite green
session-worktree-owns-stopa cross-Owns edit in worktree mode stopped the worker (blocked); the lead recovered sequentially on main
session-research-fanoutthree researcher lanes on a decomposable question, merged with contradictions and "Not searched"
session-bad-pin-garbagea corrupted model: line was refused at run open (init/checkpoint --start); zero implementer turns; target file byte-identical
session-resume-interruptedan interrupted run resumed at the first safe incomplete phase
session-worktreestwo writers in linked trees, patches applied sequentially and hash-checked

Full table and evidence pointers: E2E-REPORT.md and TEST.md in the repository. Two checks stay human-only: the guard hook's live block (needs a Desktop relaunch) and a background worker's permission denial (needs the interactive prompt).

06 / Benefits

What you get that a single context can't give you

A second pair of eyes with no memory of writing the code.

The implementer never sees your conversation; the reviewer never sees the implementer's reasoning. Independence is structural, not requested.

A bounded blast radius.

Owns paths, one writer per tree, linked worktrees for disjoint slices with hash-checked patches. A worker can edit an unowned file when it must — and the lead sees it flagged in the diff.

A record you can audit next week.

Manifest, evidence files, attest block, usage dashboard. Who ran, on which requested pin, what was verified, what it cost.

Cost routing.

Free or cheap pins for the roles that read and test, stronger pins where judgement matters. /otter:implement on the free SWE-2 pin passed every bench gate at $0.

Refusal over substitution.

A missing profile, a drifted pin, a missing evidence file or a malformed review stops the run and names the fix. Nothing silently downgrades.

07 / Caveats

The honest part

Orchestrate is an assurance instrument for work large enough that a single context degrades. It is not a way to finish small tasks faster.

It is slower and costs more than a stock lead on small work.

5-module feature, 5 trials, 2026-09-13: 6.0× the wall time and 2.9× the cost ($0.59) of a stock kimi-k3-high lead, ~21 lead turns, one dispatch at standard assurance. Two-slice runs last measured at 20–22× / 11–12× on 0.9.0. Single-worker skills are the sweet spot.

Owns is scope, not a wall, in single-tree mode.

The implementer is told to stay inside its paths and to flag anything else; only worktree mode fails closed mechanically. The lead's diff read is what catches a misbehaving model.

The dispatch pin check covers orchestrate only.

Direct /otter:implement and the other role skills are dispatched by the platform with no script seam; a corrupted model: line there may be silently substituted. Profiles run a one-command preflight.mjs self-check on spawn — defence in depth, not enforcement. /otter:models repair fixes drift either way.

Requested pins are recorded; actual worker model is inferred.

The manifest stores what was requested and the profile digest. The dashboard attributes turns by conversation tree from sessions.db and flags pin drift — that is telemetry, not a platform guarantee.

Workers don't inherit your conversation — or the web.

Every brief must be complete. Custom subagents get no web tools; the lead fetches sources into packs.

Skill bodies load at Desktop start.

Install, update or edit → full relaunch (Cmd+Q). The slash dropdown empties past /otter: — type the full command.

Two checks are still human.

Guard hook live block and background-worker permission denial are in the beta guide, not the harness.

Standard's reviewer is the lead.

That's by measurement (it caught the prototype-chain bug 5/5 vs the worker's 3/5) — but it means the second opinion at standard assurance is the same model that briefed the implementer. review=worker or assurance=high brings the clean-context worker back.