The model was never the decision. How you orchestrate the models is. This is a head-to-head benchmark of two coding harnesses on one real ticket from one of our projects, and the configuration we run now.

The premise: the workflow moves, the harness is swappable

We have written before about working human in the lead: the engineer writes the plan, agents execute it, a human reviews at the gate before anything merges. That workflow lives in our AGENTS.md files, and it has now survived two harnesses without a single change to its steps.

What we had not tested was the layer underneath. A harness decides which model reads your plan, which model writes the code, whether an implementer works in a clone or in your live checkout, and what happens when an edit lands in a file that changed underneath it. Those decisions move both the bill and the failure modes. So for the past two weeks I have been running Oh My Pi (omp on the command line) instead of Claude Code on our projects, and I ran a benchmark to turn "I like it" into a recommendation.

The test: one real ticket, six unattended runs

Synthetic benchmarks flatter tools. This was a real ticket from one of our fintech builds: CFO sign-off on unbudgeted accounts-payable routes. The scope is what a normal week produces, all of it in one ticket: routing logic, a database migration with schema mirrors, API, UI, tests, docs, and an architecture decision record.

I ran it headless through both harnesses on the same account, with Opus 5.5 as the orchestrator on both, unattended, with our AGENTS.md workflow loaded. Six runs in total. The representative pair:

  • Wall clock: 19.8 min vs 21.7 min
  • Cost, uniform Anthropic rates: $12.51 vs $14.52
  • Model requests: 212 vs 194
  • Output tokens: 167k vs 165k
  • Opus share of output: 23% vs 88%
  • Subagents: 3, all Sonnet, in isolated clones vs 3, two Opus and one Sonnet
  • Tool calls / tool errors: 306 / 14 vs 205 / 2
  • Commits: 8 vs 6
  • Build + ~11,000 tests: green on both
  • Spec checklist: 13 of 14 vs 12 of 14
  • Both runs produced the same design, down to the same new step type, the same settings key, the same migration number and the same ADR, and both finished green on the build and the full test suite. One finished two minutes faster and two dollars cheaper. The line that actually explains the money is Opus share of output.

    The results: same output, 14% cheaper

  • omp on the benchmark ticket: $12.51. Claude Code on the same ticket, same account, same orchestrator model: $14.52.
  • Opus share of omp's output tokens: 23%. Claude Code: 88%. The harness decides where the expensive model spends its tokens.
  • omp average over four clean runs: $13.40. Claude Code averaged $15.15 on the same ticket. About 14% cheaper for the same work.
  • All six runs, including the ones that did not go well:

  • cc-1: 22.2 min, $15.78, 222 requests, 6 subagents, 12/14 checklist
  • cc-2: 21.7 min, $14.52, 194 requests, 3 subagents, 12/14
  • omp-1: 24.7 min, $12.45, 199 requests, 2 subagents, 12/14
  • omp-2: 43.7 min, $34.84, 430 requests, 2 subagents, 12/14
  • omp-3: 24.2 min, $15.19, 269 requests, 4 subagents, 13/14
  • omp-4: 19.8 min, $12.51, 212 requests, 3 subagents, 13/14
  • Across the four clean runs, omp averaged $13.40 against $15.15 for Claude Code, about 14% cheaper for identical output. Two honest caveats. Claude Code was the more consistent harness run to run (21.7 and 22.2 minutes); omp ranged from 19.8 to 24.2. And omp-2 was a small disaster, at 43.7 minutes and $34.84, which is its own section below.

    Where the money goes: who writes the code

    The cost difference is not a discount. It is a different division of labour. omp's orchestrator kept Opus for planning and review and delegated tests, implementation and docs to Sonnet subagents. Claude Code did the work with Opus everywhere, which is how 88% of its output tokens ended up on the most expensive model in the account.

    Per agent, on the representative pair:

  • Orchestrator: Opus, 47 req, $4.31 vs Opus, 47 req, $4.04
  • Test author (RED): Sonnet, 91 req, $5.23 vs Opus, 49 req, $2.76
  • Implementer (GREEN): Sonnet, isolated clone, 29 req, $1.07 vs Opus, 77 req, $6.69
  • Docs / ADR: Sonnet, 45 req, $1.90 vs Sonnet, 21 req, $1.02
  • Read the whole list, not just the winners: omp's test author cost more than Claude Code's, because Sonnet needed 91 requests where Opus needed 49. The implementer line is where the switch pays for itself. Claude Code spent $6.69 of Opus on implementation; omp paid a Sonnet $1.07 to do it in an isolated clone of the repo. In our workflow that separation is native: the tests and the implementation are written in parallel, by different agents, from the same plan. The implementer never sees the tests; the orchestrator merges the test branch first, then the implementation.

    There are also levers Claude Code simply does not give us: per-agent model overrides, request budgets per subagent, capped test workers, compaction settings. Same output, and the cost still has room to move down.

    The $35 run, and what it fixed

    omp-2 belongs in this article because of what it exposed. An implementer subagent was briefed to work in a sibling worktree while its actual working directory stayed in the main checkout. Every relative path resolved against the wrong tree, its edits kept getting rejected, and it spun for most of an hour. Forty-three minutes, $34.84.

    The post-mortem went two ways. It surfaced a bug in the tool itself, which we reported upstream. And it exposed gaps in our own AGENTS.md: nothing in the workflow told an implementer that it owns its working directory, nothing told it to stop after repeated tool rejections. Both rules exist now, in every repo that runs the workflow, and they are portable to any harness. The runs after the fix are the numbers above.

    A benchmark that only records successes teaches you nothing about failure modes. The interesting property of omp-2 is that the failure was legible: rejected edits, visible retries, a transcript you can read afterwards. The agent spent most of an hour failing to apply edits, and the edits stayed unapplied. It went down spinning, not silently rewriting the wrong tree, and the fix was policy rather than luck.

    Why we stayed, beyond the 14%

    Model independence. omp talks to any provider: our Anthropic and OpenAI subscriptions, Z.ai, and our own vLLM endpoint on a RunPod GPU, all configured in one models file. Roles map to whatever model fits the budget, per project if needed. Nobody is locked to one vendor's pricing or rate limits, and the same install here currently talks to Anthropic, Z.ai and a GLM 5.3 we host ourselves. After running a local model as a serious option, that matters to us.

    Isolation as a primitive. Each implementer gets a copy-on-write clone of the checkout as its working directory and returns a branch; the orchestrator merges when it is satisfied. No hand-rolled worktrees, no agents stepping on each other in the live repo. In omp-4 the three workers ran in parallel in three clones and came back as three branches.

    Edits that fail loudly. Every file read stamps the file with a content hash. An edit must quote that hash and the line numbers it saw, and it is rejected if the file changed or the lines were never displayed. A stale edit bounces instead of silently landing three functions away from where it was aimed.

    Language-server integration. The agent queries the LSP for definitions, references and diagnostics instead of grepping, and omp runs diagnostics after every write, so type errors surface in the turn that caused them. On the benchmark repo, one references query returned all 13 call sites across four packages, on a project that had reported no language servers configured before we installed one.

    One more practical thing: it coexists with Claude Code and reads the same ~/.claude/CLAUDE.md and skills directory, so the global rules we had already written applied on day one.

    The settings we recommend

    This is close to the exact global config the benchmark ran with, trimmed to the parts that matter. It lives in ~/.omp/agent/config.yml, and one rule does most of the work: a subagent never inherits the orchestrator’s model.

  • default: anthropic/claude-fable-5-1:high (main session / orchestrator; Fable or Opus 5.5 for planning and review)
  • task: anthropic/claude-sonnet-5-5:high (routine coding, docs)
  • reviewer and security-reviewer: anthropic/claude-opus-5-5:high (the review gate)
  • scout: anthropic/claude-sonnet-5:medium (read-only exploration)
  • sonic: anthropic/claude-haiku-4-5:low (mechanical work, test runs)
  • showResolvedModelBadge: true (see which model each subagent really got)
  • isolation: enabled, merge: branch, apply: false (implementers run in copy-on-write clones and come back as branches)
  • softRequestBudget: 100 and maxRuntimeMs: 1800000 (wrap-up notice at 100 requests, hard stop at 30 minutes per subagent)
  • A few defaults are worth leaving alone, because they are doing quiet work: hash-anchored editing (described above), parallel subagent spawning from a single call, automatic context compaction, language servers shared across sessions on the same project, and long shell commands moving to the background instead of blocking the loop. We changed none of them.

    For the language servers, one install per machine covers the common stacks: npm i -g typescript-language-server for TypeScript and JavaScript, intelephense for PHP, pyright for Python. Go uses gopls, Rust rust-analyzer, Terraform terraform-ls.

    Per-project overrides go in .omp/config.yml in the repo: a cheaper default model for a small repo, a different subagent tier, or disabling a provider for a client that forbids a vendor. Arrays replace rather than extend, which keeps a project file small and explicit.

    Three rules we added, for any harness

    These came out of the omp-2 post-mortem and now live in every AGENTS.md we run, whatever the harness underneath. If you take three things from this article, take these:

  • The implementer owns its working directory. Spawn it isolated, so it gets its own clone and normal relative paths just work. If isolation is not available, the brief mandates absolute paths in every read and every edit. Never brief an agent to "work in ../some-worktree" while it sits somewhere else.
  • The tool-error brake. The same tool rejection three times in a row on the same file ends the attempt, with the exact error text reported. Never route around a rejected edit with shell tricks. An agent that cannot edit a file should say so, not find a way around the refusal.
  • One test suite at a time. A large Vitest suite forks a worker per core at roughly 3 GB each; two concurrent runs swap a 16 GB machine and leave orphan processes behind. Cap the workers, never run the API and UI suites together, and clean up forks after a cancelled run.
  • Running it headless

    The tool is a one-line install, brew install omp or npm i -g @oh-my-pi/omp, and everything is documented at omp.sh. The benchmark itself, and what a CI job or an overnight run looks like: VITEST_MAX_FORKS=3 omp -p --model anthropic/claude-opus-5-5:high --approval-mode=yolo --append-system-prompt extra-rules.md "$(cat task.md)".

    Progress lands in git and in the session transcript; each subagent has its own transcript with per-request token usage, which is where every number in this article came from.

    Adopting it on a new project

  • Start omp in the repo root and confirm the context line shows your CLAUDE.md and AGENTS.md loaded.
  • Copy in your workflow base and the three rules above.
  • Set the model roles for the project, or inherit the global file.
  • Install the language server for the stack and check that the LSP reports it configured.
  • Run one small ticket end to end, and note the cost and wall clock as the baseline before you trust it with a large one.
  • The takeaway

    The workflow was never the harness. Plan, RED tests, blind GREEN implementation, review gate: that lives in AGENTS.md and survived the switch untouched. What changed is where the tokens go. The expensive model now spends them on planning and review, implementation runs on a cheaper one inside isolated clones, and we can point the whole thing at any provider, including our own GPU.

    Measure it on your own ticket before you believe any of this, ours or anyone else's. That is what the six runs were for.

    Looking to implement agentic coding in your software development process? Contact us.