Updated 3 October 2026. This article started as my April notes on coding agents. This update corrects product capabilities against current documentation and keeps the older usage notes dated. I have not run a new six-tool benchmark for this update.
My AI agent works through tasks while I sleep. That makes finishing a task, recovering from errors, and leaving a result I can review more useful to me than a good autocomplete demo.
Which AI coding harness fits the work?
For the setup described in my original article, Claude Code was my choice for longer agent tasks. A wider shortlist makes sense today. Codex CLI runs locally and supports scripted execution. OpenCode supports project instructions and subagents. Pi can be embedded in another application, and its September releases added MCP support. Cursor also has cloud agents.
Claude Code: a starting point for a Claude-based workflow with project instructions and programmatic runs.
Codex CLI: a local coding agent with an interactive terminal and a non-interactive execution mode.
Aider: a terminal pair-programming workflow built around code edits, a repository map, and Git integration.
OpenCode: a choice for using different model providers with project rules, planning, and subagents.
Pi: an extensible agent with CLI, JSON event, RPC, and SDK integration options.
Cursor: an editor workflow plus cloud agents running in separate development environments.
These are documented capabilities, followed below by the tradeoffs I would check. A feature list alone cannot tell me which agent will finish my particular task.
What a coding harness does
The harness runs the loop around the model: send context, receive a requested tool call, execute it, return the result, and continue. File access, shell commands, editing, permissions, and context management shape what the agent can do.
It helps to separate three choices. The model does the reasoning. The harness gives it tools and manages the session. An editor or orchestration application can put several harnesses behind one interface. Changing one of these changes the system you are evaluating.
Supervised and unattended work
For supervised work, I can catch a wrong assumption while the agent is editing. For an unattended run, I need the task to have a clear stopping point, useful tests, and a result I can inspect later.
Several tools now support this second kind of work. Claude Code has programmatic execution; Codex has non-interactive mode; Cursor has cloud agents. That makes the original split between tools that need a human at every step and tools that can work alone too rigid. I would judge the whole setup, including its permissions and recovery behavior.
How much weight to give benchmarks
The original version used specific leaderboard scores and token-cost ratios to rank these tools. I am removing those comparisons because the article did not establish matching model versions, task sets, settings, and source records well enough to support that ranking.
A useful comparison needs the model, harness version, task, allowed tools, and cost conditions written down together. For my own work, I would also record whether the tests passed and how much intervention the run needed. I have not repeated that comparison for this update, so there is no new performance winner here.
Claude Code for a Claude-based workflow
My April preference came from using Claude Code for longer tasks in my own agent system. I found it better at carrying decisions through a chain of changes. That is a dated usage impression, with all the limits of one person's setup.
The current Claude Code documentation describes a coding agent that reads a codebase, edits files, and runs commands. Its programmatic interface also supports scripted runs and structured output.
Project instructions remain a practical part of that workflow. I wrote separately about how I structure CLAUDE.md after 1,000 sessions. Those instructions can reduce repeated explanation; they still need to stay accurate.
My earlier Claude Code and Codex comparison has more of the original usage context. Read it as an account of that period. Model choice, context handling, and pricing can change independently of the CLI.
Codex CLI runs on your machine
I got an architectural detail wrong in the original article. Codex CLI runs in your terminal and can inspect the local repository, edit files, and run commands there. Cloud execution is a separate way of using Codex.
It also supports non-interactive runs with codex exec. Calling it a tool that can only handle one supervised step undersells the documented execution modes.
In my April notes, I preferred Codex for contained app work and found it less consistent on longer chains. I would keep that as a usage impression from that time. It does not establish a current speed, cost, or reliability advantage for either product.
Aider vs OpenCode: which workflow do you want?
Aider is a useful candidate when I want a terminal conversation centered on editing a repository. Its repository map supplies structural context, and its Git integration can commit changes automatically. Automatic commits are configurable.
Lint and test integration needs a little precision. Aider can lint edits and fix errors. Automatic testing needs the test command and auto-test option configured. The earlier claim that every change always triggers tests was too broad.
For Aider vs OpenCode, I would start with how I want to work. Aider emphasizes the editing and Git loop. OpenCode exposes agent roles, planning, and provider configuration in the same application. Both deserve a test on the same repository before a cost comparison. The original article's token-saving ratio had too little supporting detail, so I have removed it.
OpenCode has project instructions and subagents
The earlier claim that OpenCode lacked a CLAUDE.md equivalent was wrong. Its rules documentation supports AGENTS.md and documents CLAUDE.md as a fallback when the corresponding AGENTS.md file is absent. A reader had already pointed this out in the comments.
OpenCode's agent documentation also describes primary agents and subagents, including Build and Plan workflows. These are meaningful capabilities for anyone comparing OpenCode with Claude Code.
Provider choice is still a reason to consider it. The supported connection and billing method depends on the provider. An existing subscription is worth checking against that provider's documented authentication options before assuming it pays for third-party usage.
I would compare whether OpenCode follows the repository instructions, completes the same acceptance tests, and leaves understandable changes. Those checks say more about its fit than my original description of it as only a provider switcher.
Pi vs OpenCode, and where oh-my-pi fits
Pi offers an extensible terminal agent, a JSON event stream, RPC control, and a TypeScript SDK. That makes it worth considering when a coding agent will be one component of a larger application. OpenCode offers a more predefined set of agent roles and configuration choices. Which approach fits depends on how much behavior you want to build yourself.
Pi's 0.99.0 release on 29 September 2026 added MCP, codemode, and tool search as built-in extensions. Older descriptions of Pi as having no MCP support are now out of date. That release also documents signing in with a ChatGPT subscription through the OpenAI provider. Authentication and billing need to be checked per provider.
oh-my-pi, or omp, is a separate fork of Pi with its own coding tools and workflow. Its advertised editing and subagent features should be attributed to that project. Installing Pi does not establish that someone has tested omp.
For Pi vs Codex, I would compare the integration I need. Pi documents RPC and SDK embedding; Codex CLI documents local interactive and non-interactive execution. This update is a documentation comparison. It supplies no new Pi or omp test results, and I have removed the earlier unsupported first-person Pi testing claims.
Cursor includes cloud agents
Cursor has an editor workflow and cloud agents that run in isolated virtual machines with their own development environments. The original conclusion that Cursor requires someone at the keyboard throughout a task was too broad.
For a cloud run, the environment is part of the comparison: dependencies, startup commands, repository access, and the evidence returned with the changes. For editor work, I would judge how easily I can inspect and steer the edits. I have not run a new Cursor comparison for this update.
How I would compare these on a real task
Start from the same commit with one task whose acceptance tests are written down. Record the model and tool versions, project instructions, and allowed commands. Keep separate copies of the repository so one run cannot benefit from another's edits.
Then record what happened: tests passed, missed requirements, manual interventions, elapsed time, and billed usage where available. If the products use different models or subscriptions, keep that difference in the result. It is part of the comparison.
This is a proposed evaluation method. It does not describe a new test I have already run.
What I use the comparison for now
My April notes explain why Claude Code fitted the automation I had built. Since then, I have also written about moving from terminal sessions into bb. That article covers how I manage the work around the agents.
For choosing a harness, I would narrow the list using the workflow descriptions above, then run a small task I can verify. The question I care about is whether I can understand and trust the result when I come back to it.
One thing for paid subscribers. The most relevant store product to this post is the Claude Code Prompt Pack: 50+ prompts organized by task type, pulled from real overnight sessions where I needed the harness to actually work without me. If you’re on a monthly plan, you get one free product from the store per month. That’s a good pick.
If you’re on yearly, the full store is already included. If you’re still on the free plan, this is roughly what paid unlocks in practice: the store and a weekly dispatch that goes deeper than the public posts.
I write about building with AI agents from a practitioner’s perspective. No hype, no affiliate links. Subscribe here if you want more of this.



Great article!
Though, about OpenCode, it is fully cross-compatible with CLAUDE.md (https://opencode.ai/docs/rules/#claude-code-compatibility). Reading that chapter sounds like you have skipped this completely or maybe that feature was not available back then?
Dang thats a huge insight, just what i needed right now :D im having 16gb vram and decided to cancel all my subscriptions for a while to try local models since im having 80-90t/s on qwen 3.6 35b. I want to try some of these harnesses (i finally understand what they are haha) i already tried gemini cli, codex, claude code, opencode, now im playing with PI (which is super cool, im implementing steel browser with screenshots to help me "SEE" stuff instead of be blind :D) but im really thinking about aider now...
also i was surprised you didnt mention clawbot or hermes agent, but now that i think of, they are not just harness, they are full stack so they kind a different subject here. anyway amazing article
PS: zawsze wiedzialem, ze polacy maja dryg do bycia w czolowce AI :D