The Reviewer Is Not the Gate
An AI reviewer leaves a comment: “This change can lose data when two saves race.” That might be the most valuable sentence written about the pull request. It might also be a false positive, a misunderstanding of the sto
An AI reviewer leaves a comment: “This change can lose data when two saves race.”
That might be the most valuable sentence written about the pull request. It might also be a false positive, a misunderstanding of the storage model, or a warning about a path this patch never reaches.
The comment deserves investigation. It should not, by itself, decide whether the change may merge.
Reviews and gates answer different questions
A review system helps people notice things. A merge gate enforces a rule.
An AI reviewer is useful as an advisory signal: it can scan a diff, compare patterns across files, ask whether an invariant was preserved, or suggest a missing test. Its output is probabilistic and context-dependent. A human still needs to judge whether a finding applies.
A deterministic check asks a narrower question against a defined input: did the test command succeed on this revision? Does the static rule find a forbidden import? Is the generated artifact attached to the expected source? A branch rule can then require selected checks before merge.
Neither category is perfect. Tests can omit the bug. Static rules can encode the wrong policy. Reviewers can miss a defect. The important distinction is authority: a generated comment is not automatically a trusted merge condition simply because it appears in a review UI.
A check name is not its identity
GitHub branch protection and repository rulesets can require status checks. GitHub’s documentation also cautions that users with write permission can set commit statuses. A rule that accepts any result named “Security Review” may therefore be weaker than it looks if the expected source or app is not bound.
For a required check to mean something, the policy needs to specify at least:
- which trusted application or workflow is allowed to report it;
- which commit or revision was checked;
- whether the check ran for the correct event and permissions;
- what happens when it is missing, skipped, stale, or inconclusive;
- who may bypass the rule and how that exception is recorded.
This is an identity and authority problem as much as a test problem. “Green” is a value. The repository also needs to know who produced it, for what subject, under which policy.
GitHub’s rules can be configured in different ways, and bypass privileges depend on the repository’s policy. Do not infer that a repository is protected merely because a check exists or a green badge appears. Inspect the effective rules and the trusted source of each required result.
Let AI review widen attention, not quietly become policy
There are good reasons to run model-assisted review. It can offer another pass over a large change, explain unfamiliar code, or highlight a boundary a human missed. It can also repeat itself, misread generated files, or recommend a fix that breaks a different invariant.
A useful setup keeps the AI result visible and contestable:
- The reviewer reports a finding with the affected lines and a concrete failure path.
- A maintainer validates or rejects the finding.
- Accepted findings are handled through normal review and tests.
- Deterministic checks enforce rules that are stable enough to automate.
- Human-owned policy decides whether a remaining uncertainty is acceptable.
If a team wants a model’s output to block merges, that is a governance decision. It should be explicit, versioned, scoped, observable, and reversible. The model identity, configuration, failure behavior, and bypass path become part of the policy. Otherwise an advisory tool can become a required check by accident—through a vendor default, a renamed status, or a workflow that nobody intended as authority.
A practical authority table
| Mechanism | Best at | Authority it should have |
|---|---|---|
| AI review comment | Suggesting risks and questions | Advisory unless an explicit, tested policy says otherwise |
| Unit / integration test | Checking encoded behavior | Required when relevant and trusted; limited to its covered behavior |
| Static policy check | Enforcing a crisp, repeatable rule | Blocking only when rule ownership and false-positive path are defined |
| Required status check | Providing a merge condition | Trusted only when bound to the expected source and revision |
| Human review / approval | Context, trade-offs, responsibility | Required for decisions policy assigns to a person |
This is not a ranking of intelligence. It is a map of who gets to make which decision.
The operational test
Before adding a review bot to required status checks, answer:
- What exact property does this check enforce?
- Can the result be set by a user or app outside the intended workflow?
- Does the workflow test the exact pull-request revision?
- Can untrusted pull-request code reach a privileged token or secret?
- What does a timeout or provider outage mean: fail, wait, or route to a human?
- Who can bypass, and is the bypass visible?
- What evidence would show that the policy itself is working?
GitHub’s Actions security guidance matters here: permissions should be minimized, references to third-party actions should be immutable where possible, and workflows that combine untrusted pull-request code with privileged contexts need careful separation.
The goal is not to ban AI review. It is to prevent a useful reviewer from gaining authority through ambiguity.
Capability does not appoint an owner
A model can identify a plausible bug. It cannot, by that fact alone, define the organization’s risk tolerance or accept responsibility for shipping.
A good engineering system lets models help with attention while making the decision path inspectable: a trusted check reports evidence; repository policy decides which evidence is required; a named person owns exceptional judgment.
That is the boundary Series C is about. In the next post, we will look at how to turn a pile of review comments into one focused correction cycle that can actually finish.
References
- GitHub: About status checks
- GitHub: Available rules for rulesets
- GitHub: About protected branches
- GitHub Actions security reference
Read the flagship
For the broader argument behind this boundary, read Vibe Coding Was the Prototype. Agentic Engineering Is the Next Step.. This follow-up applies its evidence-and-authority model to AI-assisted code review and merge policy.
AI assistance was used to prepare this draft. The human editor is responsible for checking current platform semantics, source freshness, and the final publication decision.
Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.