I Built a Self-Hosted AI Code Reviewer for GitHub PRs with Spring Boot
Why I built this Code review is one of the most important parts of a software development workflow, but reviewing every Pull Request manually can become repetitive. I wanted to experiment with a different approach: Wh
Why I built this
Code review is one of the most important parts of a software development workflow, but reviewing every Pull Request manually can become repetitive.
I wanted to experiment with a different approach:
What if an AI reviewer could inspect a GitHub Pull Request, identify potential problems, explain them, and post the review back to GitHub?
That led me to build CodeGuard AI.
It is a self-hosted AI code-review application built with Java 17+ and Spring Boot.
The goal wasn't to replace human code review.
Instead, the idea was to create an additional automated layer that can identify potential issues before a human reviewer spends time going through the PR.
What CodeGuard does
The workflow is essentially:
GitHub Pull Request
β
GitHub API / Webhook
β
Spring Boot Review Service
β
Fetch PR information + changes
β
AI Provider
β
Bug / Security / Performance / Best Practice Analysis
β
Quality + Risk + Confidence
β
GitHub Comments
β
Review History & Trends
The application supports multiple AI providers:
OpenAI
Google Gemini
Ollama
That makes the AI layer replaceable instead of tying the application to a single provider.
The application dashboard
The main dashboard provides two ways to start a review.
A review can be triggered manually by providing a repository and Pull Request number, or the application can work with the GitHub workflow.
It also exposes the major review categories directly in the interface:
Security Review
Performance
Best Practices
Bug Detection
The manual-review interface was particularly useful during development because it allowed me to test the complete review pipeline without depending on a live webhook every time.
Starting a review
For example, I can provide a repository and PR number and start the review.
The application then moves through the review pipeline.
The UI exposes the progress of the operation instead of making the user wonder whether the review is still running.
The basic stages are:
Reading PR
β
Finding bugs
β
Checking security
β
Generating review
This also highlights an important part of AI integrations that can easily be overlooked:
AI operations are asynchronous from the user's perspective.
A review may involve API calls, fetching repository data, preparing the PR changes, sending the analysis request, processing the response, and finally publishing the result.
What does the AI review actually produce?
After the analysis is complete, CodeGuard generates a structured review rather than returning one large block of AI text.
For one of my test Pull Requests, the result included:
Overall score: 8/10
AI confidence: 90%
Critical issues: 0
Suggestions: 2
Merge recommendation
Categorized findings
Code locations
Suggested fixes
This makes the output much easier to consume than simply displaying an LLM response.
The reviewer can immediately see the overall result and then drill down into individual findings.
Findings are categorized
Each finding is classified so that developers can focus on the type of problem they care about.
For example:
Critical
High
Medium
Low
Security
Performance
Best Practice
In one test, CodeGuard identified a potential configuration issue involving GitHub token overrides.
It also detected a repository-name case-sensitivity issue.
The important part here isn't just detecting an issue.
The system also provides:
Location β Explanation β Suggested fix
That makes the result more actionable for a developer.
Posting the result back to GitHub
One of the parts I wanted to experiment with was closing the loop.
The reviewer shouldn't have to stay inside a separate dashboard forever.
CodeGuard can post the review back to GitHub.
The dashboard therefore shows whether the review has been posted successfully.
This creates a workflow closer to:
Developer opens PR
β
CodeGuard analyzes it
β
Findings generated
β
Review posted to GitHub
β
Developer fixes issues
β
Human reviewer continues the normal review
Review history
Another feature I wanted was historical visibility.
Instead of treating every PR review as an isolated event, CodeGuard stores review results and exposes them through the review history.
This allows the application to show:
Previous Pull Requests
Quality scores
Issues found
Repository information
Review status
Review details
Review trends
The history also provides a simple trend view.
The dashboard tracks:
Quality Score vs Issues Found
over previous reviews.
This is useful because a single review score doesn't tell the whole story.
For example, if a project repeatedly receives lower-quality reviews or increasing numbers of findings, that pattern may be more interesting than any individual PR.
Making the AI provider configurable
Another engineering decision was avoiding a hard dependency on one AI provider.
CodeGuard supports:
OpenAI
Gemini
Ollama
The provider can be selected/configured through the application settings.
This also makes the project interesting for developers experimenting with local models.
For example, a developer can use a hosted provider during development and experiment with a local Ollama setup when self-hosting is preferred.
Custom review instructions
The settings also allow custom review instructions.
For example, a team could specify additional standards such as:
Flag public methods without Javadoc.
Flag endpoints that don't validate input.
Flag direct System.out usage instead of logging.
Flag newly thrown generic exceptions.
These instructions are added to the review process alongside the built-in review criteria.
This makes the reviewer more adaptable to a project's own engineering standards rather than relying only on generic AI code-review prompts.
The Spring Boot architecture
The backend is built around Spring Boot.
The main pieces include:
GitHub Integration
β
PR / Diff Retrieval
β
Review Service
β
AI Provider Layer
β
Structured Review Result
β
Finding Classification
β
GitHub Comment / Dashboard
β
Review History
The project uses:
Java 17+
Spring Boot
Spring Security
GitHub REST API
AI provider integrations
REST APIs
Maven
HTML/CSS/Vanilla JavaScript
The important architectural decision was keeping the AI provider layer replaceable.
That means the application isn't fundamentally tied to one model provider.
What I learned building it
- AI output needs structure
Sending code to an LLM is relatively easy.
Building a useful code-review product around that response is considerably more involved.
You need to think about:
What constitutes a finding?
How should severity be represented?
How do you identify the relevant code location?
How do you avoid overwhelming the developer?
How should findings be presented?
How should the result get back into the existing GitHub workflow?
- Confidence and severity aren't the same thing
An AI system can be highly confident about a low-severity issue.
So I wanted CodeGuard to keep concepts such as:
severity
and
AI confidence
separate.
That gives developers more information when deciding which findings deserve attention.
- Human review still matters
The purpose of CodeGuard isn't:
"AI reviewed the PR, therefore merge it."
The more useful model is:
AI performs an additional automated review layer, while humans remain responsible for the final engineering decision.
That distinction became important as I designed the workflow.
What could be improved next?
There are several directions I'd like to explore:
Better inline GitHub comments
More precise code-location mapping
Improved false-positive handling
Repository-level review policies
Custom severity rules
More granular review configuration
Better webhook lifecycle handling
Additional local-model support
More detailed review analytics
The project is intentionally designed as a self-hosted foundation that developers can customize, rather than a hosted SaaS product.
Final thoughts
Building CodeGuard taught me that an AI code reviewer is not simply:
GitHub API + LLM = code reviewer.
The interesting engineering work happens around the model:
GitHub integration β PR processing β review orchestration β structured findings β severity β confidence β developer feedback β history.
That's where the application becomes an actual workflow rather than just an AI prompt.
Want to explore the implementation?
I built CodeGuard AI as a self-hosted Spring Boot source-code project for developers who want to study, customize, and build on this type of AI code-review workflow.
Get the CodeGuard AI source code on Gumroad: https://javacoder716.gumroad.com/l/codeguard-ai
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.






