AI AgentsIntroducing Critic, a Review Agent Running on Another Company’s Model (Codex)
· AlgoSu
- #multi-model
- #codex
- #critic
- #harness
One thing up front: all twelve agents run on Anthropic's Claude models.
Code review was no different. An agent for security and contracts and the lead agent for planning and decision records split the reviews between them, but both were from the same Claude model family. The same model misses the same things, so they couldn't catch each other's blind spots.
Around then, OpenAI released a tool that let you call its coding model, Codex, from inside Claude Code. That opened a way to bring in the eyes of a different model family.
Problem
The harness (the rules, tools, and runtime the agents work in) was tied to one company's line of models (a single model family), and code review was handled entirely by that family too, so nothing could catch the blind spots they shared. It needed another perspective.
Decision
Put Critic, a review agent running on Codex, in the seat right before merge. Not a replacement for the existing agents, but one added seat. One that doubts the consensus instead of agreeing with it.
Result
Eighteen rounds of cross-review caught 8 critical issues (P1) and 9 recommended fixes (P2). But the Critic missed process violations, and a bigger bias surfaced: the harness itself was bound to one model family. The next goal is a harness where models swap like parts.
Why Add, Not Replace
I never considered a full replacement.
The twelve agents had already settled into their roles. Each had a stable role definition, reference documents, and ways of reaching the others. There was no reason to tear that down. And switching wholesale from one model family to another wasn't guaranteed to be the answer. That would only change what I depended on.
So I decided to add. Keep the twelve as they were, and create one new seat for a different model family.
Why Right Before Merge
There was a reason the new seat went into the review stage.
I wanted one last look at correctness, security, concurrency, and whether a change could be rolled back, just before merge. That was also where a different perspective was worth the most. When models from the same family review each other, they share the same blind spots from the same training data. If they agreed, the code passed.
Bring in a different family right before merge, and you get a seat that questions that agreement. A seat that breaks consensus. That's where the new agent belonged.
That's why I named it Critic. Not someone who agrees, but someone who doubts.
18 Rounds of Another View
I first saw the Critic really at work in sprint 135 (a sprint is a short work cycle). AlgoSu has run well over a hundred sprints; this story is from around the 130s. I merged five PRs in a row and called the Critic on each. One review-and-fix cycle counts as a round. Here's how it went.
3rounds
Work batch A (PR #167)
AI usage-limit check prototype (aiQuotaCheck PoC)
4rounds
Work batch B (PR #168)
4 critical / 2 recommended
3rounds
Policy sync (PR #169)
Aligning the error-handling filter (errorFilter) policy
5rounds
Work batch C (PR #170)
2 critical / 3 recommended
3rounds
Work batch D (PR #171)
1 critical / 1 recommended
In total, 18 rounds found and fixed 8 critical issues and 9 recommended fixes before merge. Only three of the PRs above have counts listed, so their sum (7 critical, 6 recommended) is lower than the total; the counts for the other two PRs (work batch A and the policy sync) weren't recorded separately.
What interested me was the kind of defect. Most of the critical issues were in code that the same model family would have agreed on and passed. A module that missed side effects because all the attention went to the integration. A branch that leaned on a heuristic and still had a race condition. A filter that mixed up business errors and infrastructure errors.
Each piece of code worked. It just had never been seriously challenged. Codex was where that challenge came in. Once a model that didn't share the same training distribution joined, the consensus started getting questioned.
That's when I first felt the value of "model diversity." I didn't need a smarter model. I needed a model that had learned differently.
The Critic's Limits
The joy didn't last.
After sprint 135, while tidying up the notes, I realized something: the Critic had been there all along, and it still didn't stop the direct push to main.
In the sprint just before, an agent edited a few lines of a work-record document, then committed and pushed straight to the main branch, with no branch and no PR. The edit was four lines of a decision record (ADR): it changed the status to completed and added two or three lines of cleanup.
CLAUDE.md, the rules file the agents load automatically every session, says it in bold: "Agent branch discipline: direct push to main is strictly forbidden." The rule had been tightened after an earlier violation and repeated more than once. In the notes for the next sprint, the agent left a single line about itself: "Self-recorded policy violation."
It wasn't the first time. A few days earlier, it had concluded from a shell glob (a wildcard pattern) that a directory didn't exist. I corrected it. The directory was there. A few days later, it made the same mistake with the same pattern.
The agents were ignoring the agreed workflow more and more often. Even with the Critic in place, the same model kept making the same mistakes.
The Critic checks whether code changes are correct, right before merge. It doesn't catch process violations. Committing to main without a new branch isn't broken code. It's skipping the workflow. No PR is created, so the Critic is never called.
The blind spots in code review shrank, but process bypasses remained a separate problem.
Where the Critic sits, and its blind spot
The Harness Was Biased Too
That pushed me one step further back.
Looking at the harness as it stood, everything from directory names to tool names was built around one model family's environment (for example, the CLAUDE.md rules file and the .claude/ settings folder). Seating Codex in one more spot didn't give the system's skeleton real model diversity. I had put a different model into a single review seat, but the skeleton was still bound to one model family.
That was the real problem.
This was similar. What I depended on had moved from an external service to an AI model, but the pattern was the same. When a system is bound to one family, it shakes when that family shakes.
The same lesson had come back at a different layer.
Toward a Model-Agnostic Harness
Which model is best changes over time. Codex may lead at coding today; next quarter it could be Claude, and after that Gemini. There's no way to know in advance, and no need to.
What matters is being able to swap the moment that happens. That's what freedom from dependency means.
The next step has two layers.
First, a per-agent model switch. The twelve role definitions stay as they are; which model family fills each seat becomes a setting. The Critic seat could hold Codex now and a different model next quarter.
Further out, freeing the harness itself from any one model or runtime. Tools and environments become swappable parts, and the system's skeleton sits on top of them, so a new model can be plugged in the moment it arrives, without waiting.
The start was small. One seat for a different family, 18 rounds of review, eight critical issues caught. I don't plan to stop there. I'm taking the lesson from when Baekjoon vanished and applying it again, more deeply this time, at the model layer.
A system freed from dependency lasts longer.