I Taught a Computer to Spot Fake Websites With 96% Accuracy
What if a computer could tell a fake website is fake just by looking at its URL — without ever opening it? Turns out, yes. Here's how I built my first real machine learning project. The Problem: Phishing Is Eve
What if a computer could tell a fake website is fake just by looking at its URL — without ever opening it? Turns out, yes. Here's how I built my first real machine learning project.
The Problem: Phishing Is Everywhere
If you've ever gotten a message saying "Your bank account is suspended, click here now," you've seen phishing in action — fake websites disguised as real ones, built to steal your password or credit card number.
Traditional defenses (blacklists of known bad URLs) are always playing catch-up. Attackers spin up new fake domains faster than blacklists can be updated.
The Idea: Teach the Pattern, Not the List
Instead of memorizing bad URLs, what if a model learned the shared traits of phishing sites? Things like:
- Does the site have a valid SSL certificate?
- Do the page's internal links point somewhere suspicious?
- Is the domain brand new or well-established?
The Experiment
I used the UCI Phishing Websites dataset — 11,055 real websites, each labeled phishing or legitimate, described by 30 features.
I trained and compared three classic ML algorithms:
| Algorithm | Accuracy |
|---|---|
| Logistic Regression | 92.45% |
| Random Forest | 96.70% 🏆 |
| SVM | 94.71% |
The Most Surprising Part
When I checked which features mattered most, SSL certificate state and anchor link behavior dominated — by a wide margin over the other 28 features.
Why? Because phishing sites usually:
- Can't get a valid SSL certificate for a fake domain
- Copy the real site's design, so internal links accidentally point back to the real domain
The model figured this out on its own — I never told it to focus on SSL.
The Real Takeaway
The biggest lesson: you don't need to be an expert to start. The dataset was ready-made, the tools (Python + scikit-learn) are free, and the steps are well-documented. What it actually took was patience and consistency.
Full code and details are on GitHub:
🔗 https://github.com/eln2mac-has/phishing-detection-ml
If you try something similar or have questions, drop a comment below 👇
Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.