AI AgentsWhen to Bring In a High-End AI Model (Fable 5)
· AlgoSu
- #multi-model
- #model-selection
- #refactoring
- #cost
One thing up front: all twelve agents run on Anthropic's Claude models. Day to day, I used Claude Opus (Anthropic's top model until then) and Claude Sonnet (a lighter one).
The session-limit notice showed up in the last sprint of a five-sprint cleanup (a sprint is a short work cycle).
The auth & security agent (Gatekeeper) was in the middle of pinning eight third-party GitHub Actions, used in sixteen places, to exact commits (SHAs). The run log (the .out file) held two lines: "session limit · resets 7:10pm." That is the notice that Claude usage has hit the cap allowed for a set period. The agent had stopped mid-task, and its uncommitted changes were still sitting there. My reaction was less surprise than "Oh, right. The pricier model this time."
Fable 5, Anthropic's high-end model a step above Opus, came out on June 10, 2026, and I switched that same day. The first sprint instruction I gave it was: "Analyze every security issue and improvement point in the codebase, then build a sprint plan."
Problem
A more capable model was available. But whether it was always the right choice was a different question.
Decision
Use Claude Opus and Sonnet day to day, and bring in Fable 5, the newer model above them (the premium model from here on), only for large cleanups.
Result
What six sprints (one audit plus five cleanups) taught me wasn't a model's limits. It was the importance of timing.
The Launch-Day Full Audit
Giving Fable 5 a full codebase audit as its first job was, in hindsight, the right kind of test. It wasn't "let's see if it's better." It was closer to "let's give it the hardest thing we have."
The result was a decision record (ADR), ADR-030. In this post, ADR-030 means both that audit and the cleanup plan that came out of it.
It scanned three areas in parallel (security exposure, code quality, and CI and infrastructure), and every suspicious conclusion was checked by opening the code directly. That filtered out five wrong calls. For example, it reported that the rollback logic for failures (Saga compensation: when a multi-service operation fails partway, undo the earlier steps) was missing, but it was there: compensateGitHubFailed/compensateAiFailed. It also reported that the code editor (Monaco) wasn't lazy-loaded, but next/dynamic already handled that. Separating the broad sweep from the careful check is what caught these mistakes.
The final verdict was zero high-risk findings, confirmation that the fundamentals were solid. Three medium, five low, seven improvements. From that came a five-sprint plan to work through them. One audit sprint plus five cleanup sprints: those are the six sprints this post talks about.
Five Sprints In
Sprint 1 · Quick security wins
One place to mark public APIs · Input validation for the events API · Isolate outside input that ends up in prompts · Request-header cleanup
Sprint 2 · Operations runbooks
Encryption-key rotation procedure · Procedure for reprocessing failed messages (docs only)
Sprint 3 · Backend decomposition
study.service 823 lines → 5 services · saga-orchestrator 516 lines → 3 modules
Sprint 4 · Frontend decomposition
AddProblemModal 805 lines → 4 files · settings 844 → 252 lines · Redis logging cleanup
Sprint 5 · Supply chain · CSP · CI
Pin 8 external GitHub Actions to exact versions · Move scripts embedded in CI into files · CSP (Content Security Policy) experiment (deferred)
The plan went as expected. But what stood out more than performance was cost.
Token usage was clearly different. Deeper reasoning means the model holds more context in one pass, and that means burning more tokens. In the fifth sprint, the security agent hit the session limit partway through the version-pinning work, a first. It was the first session interruption since switching to Fable 5.
Regressions at the Seams
Across the backend and frontend decomposition sprints, one pattern kept repeating.
When you split large files apart, the regressions (things that used to work and now don't) gather at the new boundaries.
When study.service.ts was split into five files, what broke wasn't the logic inside any one file. It was the boundaries between them. The redis.keys O(N) problem, which scans every Redis key, had been in the code before the split; the split is what finally brought it into view of Critic, the review agent that checks code right before merge.
The frontend decomposition went the same way. First, the review stages: the review agent, Critic, looks at code in two stages. ① Automated review (Auto-Critic) checks each change as it's committed. ② The final pre-merge review looks at the whole change against the base branch, not a single commit, and runs in rounds 1, 2, … until it's clean. Here's where the frontend problems were caught:
- Text not wrapped in translation keys → ① automated review
- A dropped
setError(null)reset → ② final review, round 1 - A missed derived value (one computed from other state) → ② final review, round 2
Round 3 of the final review was the first clean pass.
These decomposition sprints ran on Fable 5 too. Even so, these problems were only caught after rounds of review. Here is how much review the ADR-030 five-sprint plan took:
15rounds
Critic review rounds (two decomposition sprints)
Backend 7 + frontend 8 (including final reviews)
6findings
Regressions at boundaries
Found by automated review and final review rounds 1–2
Every boundary regression only surfaced at the review stage: in the automated review or the final review of the whole change. Even with the premium model doing the work, what caught these regressions was a second look from a different model.
I haven't measured whether a premium model cuts the number of review rounds. I never ran the same decomposition on the normal models, so there's nothing to compare against.
When a Premium Model Pays Off
The heart of the strategy is choosing where to use it.
Regular sprints (new features, bug fixes, documentation updates) are fine on Opus and Sonnet. When you're adding code on top of established patterns, the difference in reasoning depth barely shows. Cost efficiency matters far more.
A premium model earns its return on investment (ROI) in other cases:
01
Work where a normal model would go through several rounds of final review
Large decompositions with many boundaries, security audits, refactoring that cuts across modules. Fewer review rounds translate directly into lower cost.
02
Sprints that pay down the principal of technical debt
Like ADR-030, a concentrated effort to clear out problems that built up over time. Spending once here raises the quality of the normal-model sprints that follow.
03
Unfamiliar territory
A full codebase audit from scratch, or designing a new area. The less context you start with, the more reasoning depth matters.
When to reach for the premium model
That's the core lesson of these six sprints. Spending on a premium model isn't for that one sprint. It's to raise the baseline for many sprints after it. Like paying down debt: retire the principal, and the interest (review-round cost) shrinks.
Where This Strategy Fails
Reading this as "always use the premium model" would be wrong.
It rests on a few conditions.
First, wholesale refactoring has quickly diminishing returns. The ADR-030 decomposition worked because there was clear "principal" to pay down: 823 lines, 516 lines. Splitting code that's already well structured adds regression risk for little gain. Use it where the pain is obvious, not everywhere.
Second, bigger doesn't automatically mean more accurate. The decomposition sprints ran on Fable 5 as well, yet the regressions at the split boundaries still only showed up at review. Even with a stronger model, a second look from a different model was still needed.
Third, the cost is real. Hitting the session limit in the fifth sprint was a warning. Run every sprint on a premium model and token usage, the risk of interrupted sessions, and cost all rise together. That's why it has to stay selective.
The Previous Post and This One
In my earlier post, “Introducing Critic, a Review Agent Running on Another Company’s Model (Codex)”, I wrote about a system that isn't tied to any one model, bringing in Critic, the review agent, as an addition rather than a replacement, and moving toward a harness (the rules, tools, and runtime the agents work in) where any model can take any seat.
This post is the sequel. It answers the next question that comes from moving toward that freedom: "So which model, when, and where?"
If you're locked into one model, there's nothing to choose, so the question doesn't even exist. Only by moving toward a setup where models can be swapped does a strategy become possible, and timing becomes something you can choose.
The day Fable 5 launched, I put the brand-new model on a full codebase audit. The results were good. But what mattered more came after: learning when to reach for it again, and seeing that judgment settle into a frame: not a performance comparison, but when to pay down technical debt.
When the next large refactoring is needed, I'll bring the premium model out again. Not now. The boundaries are already clean, and regular sprints are running well on top of them.