When Human in the Loop Becomes Theatre
I put Claude Code and Codex in a pipeline that plans, implements, and cross-reviews changes. The most revealing part was how quickly I started trusting it.

After several clean pull requests in a row, I caught myself approving faster than I was reading.
That was the most revealing moment in an experiment I've been running on my own machine. I had put Claude Code and Codex into a pipeline where they plan changes, debate the approach, implement, and review each other's work. I wanted to understand how much useful work I could delegate, and whether having two models challenge each other was worth the cost.
The workflow worked well enough that my behaviour changed. I was still there. I still had the final approval. But I was starting to treat that approval as confirmation of what I expected, rather than a decision I needed to earn by checking the work.
That's where “human in the loop” gets uncomfortable. A person can remain in the process while their attention quietly leaves it.
The experiment: two models, one requirement#
My setup is deliberately more than a prompt running on autopilot.
It starts with a requirement. One model turns that requirement into a full ticket, grounding the acceptance criteria in the existing codebase and business logic. It looks for the most pragmatic implementation path before anyone starts changing files.
Then Claude Code and Codex debate the plan. They implement the change and cross-review each other's work. When they disagree, one is assigned to make the final call.
I watch the exchange. I'm interested in where they catch each other's mistakes, where they simply agree, and whether the reasoning behind a decision holds up. I also track the token cost of each stage.
From the inside, it feels like supervising a small, very fast team that argues in text.
The question I'm testing is whether that extra conversation buys enough quality to justify its price. A longer discussion is easy to produce. A useful challenge to a bad assumption is much more valuable.
Agreement needs evidence#
Cross-review is appealing because it gives the first implementation another chance to be challenged. A second model can question the scope, notice a missed case, or suggest a simpler approach.
But two models agreeing doesn't settle whether a change is correct.
They can both accept the same incomplete requirement. They can both work from the same mistaken assumption about the codebase. They can review whether the implementation follows the plan without questioning whether the plan solves the right problem.
Assigning one model the final call resolves the disagreement so work can continue. It doesn't make that decision true.
That's why I care about what a review actually establishes. “Looks good” tells me very little. A review that identifies an assumption, points to the relevant code, and explains how the behaviour was checked gives me something I can assess.
The useful output of the debate is a decision with reasons and evidence attached.
Success changes the reviewer#
The benefits are real. Delegating routine implementation creates more room for architecture, product thinking, and decisions that need judgment.
But a run of good results also makes it tempting to spend less attention on the next one. I felt that in my own review. Each clean PR made the following approval feel a little more obvious.
The interface doesn't reveal that shift. There's still a human reviewer and an approval button. The process can look exactly the same while the check becomes weaker.
For me, a useful warning sign is whether I can explain what I verified before approving. If all I can say is that the agents agreed and the checks passed, I need to understand what those checks covered and what they left open.
A passing test suite is evidence about the behaviours it tests. It can't answer a product question nobody translated into a test, or protect an existing behaviour nobody thought to preserve.
Define good before the code exists#
This is where I think the human role needs to move. More of the work belongs at the beginning, when the requirement and acceptance criteria are still being shaped.
A ticket that says “add filtering” leaves a lot for the agents to decide. Which records should a user be able to filter? What happens when nothing matches? Does filtering happen before pagination? Which existing permissions must still apply?
Those are product and engineering decisions. If they're left implicit, a convincing implementation can still deliver the wrong behaviour.
Before delegating a change, I'd want the requirement to establish:
- The user outcome the change should deliver.
- The existing behaviours it must preserve.
- The constraints that limit the solution.
- The evidence needed to call it complete.
- The uncertainties that need a human decision.
That gives both the implementer and the reviewer a shared target. It also makes it easier to spot a polished solution to a different problem.
Make review capable of changing the outcome#
A meaningful human check needs a chance to stop or redirect the work. Otherwise, it's ceremony.
For an agent-generated PR, I want to be able to answer a few concrete questions: What changed for the user? Which assumptions did the implementation make? What was actually tested? What remains uncertain?
The depth of that review should depend on the change. A copy edit and a permissions change deserve different attention. A change that touches access control needs evidence that disallowed actions still fail, as well as evidence that the intended path works.
I also want the cross-review to challenge the implementation against the original requirement. Reviewing only the author's explanation makes it too easy to inherit the author's framing.
The practical test is simple: could this review uncover a reason to reject the change? If the process only produces another confident summary, I haven't gained much assurance.
The conversation has to earn its cost#
Tracking tokens matters because each additional planning or review stage has a price. The experiment needs to tell me which stages are useful enough to keep.
A review that catches a business-logic mistake before merge may be worth a substantial discussion. Several rounds that restate the same conclusions may add cost without improving the result.
Alongside token usage, I'd want to track what each stage contributes: meaningful issues found, changes made because of review, human rework still needed, and problems discovered later.
I don't yet have a universal answer for how much autonomy is too much. That's what makes the experiment useful. It exposes the trade-off in actual work, including the way my own attention responds to repeated success.
The loop is moving#
I still want to delegate more implementation. I also want to get better at defining what good means and designing checks that catch a confidently wrong change.
That takes engineering judgment before the work starts, evidence while it runs, and attention at the point where someone accepts responsibility for the result.
The moment I started approving faster than I was reading gave me a better question to ask about the whole pipeline:
What would make me reject the next PR, and would my process actually show it to me?If I can't answer that, my presence in the loop is doing less than I think.
Adapted and expanded from my experiment featured in Adroiti Technologies' LinkedIn post.


