Automated Code Review, Service Tiers and Multiple Parallel Agents - Weekly Reflections #2
There’s a lot of talk about pull requests being reviewed and deployed automatically, with no human in the loop. This week I tried to answer the question that idea leaves open for me: which PRs? Which PRs can sk
There’s a lot of talk about pull requests being reviewed and deployed automatically, with no human in the loop. This week I tried to answer the question that idea leaves open for me: which PRs?
Which PRs can skip human review
The answer depends on two things: what guardrails a project already has, and how critical it is.
Before even considering fully automated review, I’d expect a project to have at least:
- a type checker, for dynamically typed languages;
- linters and formatters;
- checks for cognitive and cyclomatic complexity;
- a good test pyramid: unit, integration, component, smoke and end-to-end tests;
- a combination of static tools and agent skills that keeps a minimum of organization in the code;
- a CI/CD pipeline that works and is fast;
Which pages still need an editor?
Even with all of that, projects differ. To think about criticality, I borrowed the idea of service tiers from OpsLevel, a service catalog platform:
- Tier 0: business-critical. Downtime can stop essential operations. Needs high availability and fast recovery;
- Tier 1: downtime affects important business flows.
- Tier 2: degradation affects secondary features. Longer recovery periods are tolerable;
- Tier 3: internal services, auxiliary tools or non-critical workloads. Lower availability and support requirements;
Tiers 0 and 1 are clearly out of the question for 100% automated review and deploy. Tiers 2, and especially 3, can be handled that way without major consequences. Something broke? Roll back, fix it and release it again.
OpsLevel’s tiers only describe availability, though. A bug doesn’t have to take a service down to cause a disaster. So I defined a second classification, for the impact of errors:
- Tier 0: errors can cause significant financial loss, expose sensitive data or violate regulations;
- Tier 1: errors affect important business flows, but the impact is limited or recoverable;
- Tier 2: errors have low impact and can be fixed without relevant consequences;
Given the basic guardrails above, my view is that services in OpsLevel Tiers 2 and 3 that are also Error Tier 2 could have review fully automated. The others, no.
I like these classifications because they make me ask concrete questions. For example, what would make me comfortable with a fully automated review and release for an OpsLevel Tier 1 service? Performance tests? Performance benchmarks? Concurrency tests?
All of this gets much harder in large services or monoliths. When everything is split into separate repositories, it’s far easier to apply the tiers and decide what can be reviewed and deployed automatically. Inside a monolith, especially a messy one with no modules organized around the business (which is the rule), it’s a very different conversation.
Modern monoliths are like this house of cards: somehow, it holds.
One more guardrail applies even to small services: checking that existing tests unrelated to the change weren’t modified. Without it, a change can pass CI because the test that would have caught it was edited. Maybe the code of test cases should always be reviewed by a person.
Paying down debt with parallel agents
This week I upgraded several services to Django’s long-term support releases. One went two steps, from 3.2 to 4.2 and then to 5.2. Another went straight to 5.2.
While doing the upgrades, I found a gap in our test suite and started closing it by adding component tests to the services I was touching. Doing both at the same time proved pretty manageable with tmux, several Claude Code instances and git worktrees. Each worktree is a separate checkout of the same repository on its own branch, so every agent had its own folder to work in, and tmux kept all the sessions open side by side in one terminal. While one instance worked on an upgrade, another wrote tests, and I moved between them reviewing.
For me, that’s a very good example of how AI is an extraordinary tool for paying down technical debt and keeping the lights on.
New releases of these services can now rely on the new tests, which removes part of the testing bottleneck.
By the way, I didn't have much context one some of those services, so building the tests was a very effective and productive way to learning how things work under the hood.
Parallel effort, shared payoff.
Still wondering
- What would make a fully automated review and release safe for a Tier 1 service? Performance tests, benchmarks, concurrency tests, or something else? Canary release + feature switches and or flags?
- How do you apply tiers inside a monolith that has no business-oriented modules?
- Should changes to test code always get human review, even when everything else is automated?
Let's see what the future holds.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.

