Turning an SEO audit into tasks your AI coding agent can actually fix (Claude Code, Cursor, Codex)
A typical SEO audit export is a spreadsheet with a few hundred rows: "Missing canonical (18 pages)," "Duplicate title (11 pages)," "Orphan page," "Redirect chain." Pasting that into Claude Code, Cursor, or Codex with "fi
A typical SEO audit export is a spreadsheet with a few hundred rows: "Missing canonical (18 pages)," "Duplicate title (11 pages)," "Orphan page," "Redirect chain." Pasting that into Claude Code, Cursor, or Codex with "fix these" usually goes badly. The agent edits the wrong template, "fixes" a warning that didn't matter, or changes robots.txt in a way nobody reviews.
Coding agents are good at SEO fixes because most technical SEO problems are template problems, and a template is code. The trick is the translation step between the audit and the agent. This post is the process I use.
Step 1: Sort findings into three buckets
Not every SEO finding is a coding task. Before anything goes to an agent, I sort the list.
Delegate to the agent (deterministic, lives in code, easy to verify):
- Missing, duplicate, or self-contradicting
rel="canonical"tags - Duplicate or empty
<title>and meta descriptions produced by a template - Leftover
noindex(meta tag orX-Robots-Tagheader) on production routes - Sitemap problems: wrong host, non-canonical URLs, URLs that 404 or redirect
- Internal links pointing at redirects or 404s
- Redirect chains (A→B→C) that should be a single hop
- Missing or invalid JSON-LD (
Organization,SoftwareApplication,BlogPosting,BreadcrumbList) - Missing
altattributes on content images, and heading-level order in components
Do it with the agent, but decide yourself (code changes that encode a business decision):
-
robots.txtrules, especially for AI crawlers - Which URL is canonical when two pages overlap
- Redirect maps after a URL restructure
-
hreflangsets
Keep away from the agent (not code problems):
- Content quality, search intent, and topic gaps
- Backlinks and authority
- Anything that's really a keyword-strategy decision
The first bucket usually covers most of what a technical audit flags. The second bucket is where an agent can do real damage if you let it decide on its own.
Step 2: Group by cause, not by URL
Audits report symptoms per URL. Agents fix causes. "Missing canonical on 18 pages" isn't 18 tasks. It's usually one layout or one generateMetadata-style function that never sets a canonical. Before writing tasks, look at the affected URLs and ask what they share: a route pattern (/blog/[slug]), a layout, a CMS content type.
One task per cause keeps diffs small and reviewable, and it stops the agent from patching 18 pages one by one, which you'd then have to maintain.
Step 3: Write each task so it can be verified
This template has worked well for me with all three agents:
## Task: Add self-referencing canonical to blog posts
**Finding:** 18 URLs under /blog/* have no <link rel="canonical">.
**Evidence:** `curl -s https://example.com/blog/hello-world | grep -i canonical` returns nothing.
**Likely cause:** app/blog/[slug]/page.tsx builds metadata without `alternates.canonical`.
**Change:** Set the canonical to the absolute production URL of the post
(https://example.com/blog/<slug>), built from the site's configured base URL.
**Scope:** Only files under app/blog/. Do not touch robots.txt, redirects, or other routes.
**Acceptance criteria:**
- Every /blog/<slug> page renders exactly one canonical tag.
- The canonical is absolute, uses https, the production host, and no query string.
- No other route's <head> output changes.
**Verify:** `npm run build`, then run scripts/check_seo.py against urls-blog.txt.
What each part does:
- Evidence stops the agent from "fixing" something that isn't broken. If it can't reproduce the evidence, it should say so.
- Likely cause is a hint, not an order. Agents are good at confirming or correcting it once they open the file.
- Scope is the most important line. Without it, agents tidy up nearby code, and that's how a canonical fix turns into a 40-file diff.
- Acceptance criteria are what you'll review against. Make them observable in rendered HTML, not in source code.
- Verify gives the agent a way to check its own work before handing it back.
Step 4: Give the repo standing instructions
Rules that apply to every SEO task belong in the repository, not in each prompt. Codex and Cursor read an AGENTS.md file in the repo. Claude Code reads CLAUDE.md, and you can pull in a shared file from there with an @AGENTS.md import line. I keep the SEO rules in one place:
## SEO rules
- Production base URL: https://example.com (never localhost, preview, or staging hosts).
- Canonicals: absolute, https, production host, no query strings, self-referencing by default.
- Never add `noindex` or change robots.txt without an explicit instruction in the task.
- Titles come from the page's own data; never hardcode the site name as the full title.
- JSON-LD must only state facts visible on the page (no prices or ratings that aren't shown).
- After any <head> change, run: python3 scripts/check_seo.py urls.txt
That last rule matters most. It turns "looks right" into a command with an exit code.
Step 5: Verify against rendered HTML, not the diff
SEO bugs live in the HTML a crawler receives. A diff can look perfect while a parent layout overrides the canonical, or a client component injects a second <title>. So the check should fetch real pages. Here's a small standard-library Python script I give agents (and use myself). It takes a file of URLs, fails on missing or duplicate canonicals, canonicals pointing elsewhere, noindex, redirects, error status codes, and titles duplicated across pages:
#!/usr/bin/env python3
"""check_seo.py: verify basic SEO invariants for a list of URLs (one per line)."""
import sys
import urllib.error
import urllib.request
from urllib.parse import urlsplit, urlunsplit
from collections import defaultdict
from html.parser import HTMLParser
class HeadParser(HTMLParser):
def __init__(self):
super().__init__()
self.titles, self.canonicals, self.robots = [], [], []
self.in_title = self.in_body = False
def handle_starttag(self, tag, attrs):
a = {k: (v or "") for k, v in attrs}
if tag == "body":
self.in_body = True
elif tag == "title" and not self.in_body: # ignore <title> inside inline SVGs
self.in_title = True
self.titles.append("")
elif tag == "link" and "canonical" in a.get("rel", "").lower().split():
self.canonicals.append(a.get("href", ""))
elif tag == "meta" and a.get("name", "").lower() == "robots":
self.robots.append(a.get("content", ""))
def handle_endtag(self, tag):
if tag == "title":
self.in_title = False
def handle_data(self, data):
if self.in_title:
self.titles[-1] += data
def norm(u):
"""Treat https://a.com and https://a.com/ as the same URL."""
parts = urlsplit(u)
return urlunsplit(parts._replace(path=parts.path or "/"))
def check(url):
req = urllib.request.Request(url, headers={"User-Agent": "seo-check/1.0"})
try:
with urllib.request.urlopen(req, timeout=15) as r:
final, x_robots = r.geturl(), r.headers.get("X-Robots-Tag", "")
html = r.read().decode("utf-8", "replace")
except urllib.error.HTTPError as e:
return None, [f"HTTP {e.code}"]
p = HeadParser()
p.feed(html)
problems = []
if norm(final) != norm(url):
problems.append(f"redirects to {final}")
if len(p.titles) != 1 or not p.titles[0].strip():
problems.append(f"{len(p.titles)} <title> tags")
if len(p.canonicals) != 1:
problems.append(f"{len(p.canonicals)} canonical tags")
elif norm(p.canonicals[0]) != norm(url):
problems.append(f"canonical -> {p.canonicals[0]}")
if any("noindex" in v.lower() for v in p.robots + [x_robots]):
problems.append("noindex")
title = p.titles[0].strip() if p.titles else None
return title, problems
if __name__ == "__main__":
urls = [line.strip() for line in open(sys.argv[1]) if line.strip()]
by_title, failed = defaultdict(list), False
for url in urls:
title, problems = check(url)
if title:
by_title[title].append(url)
print(("FAIL " if problems else "ok ") + url + (" " + "; ".join(problems) if problems else ""))
failed |= bool(problems)
for title, pages in by_title.items():
if len(pages) > 1:
print(f"FAIL duplicate title {title!r} on {len(pages)} URLs")
failed = True
sys.exit(1 if failed else 0)
Run it against a local production build (npm run build && npm start) with a URL list that covers one page per template, then again against production after deploy. List URLs in their canonical form. A trailing-slash mismatch counts as a failure on purpose, because it's a real inconsistency.
It's deliberately narrow. It doesn't judge title wording or content. It checks the invariants an agent is most likely to break.
Step 6: Review like it's a migration
Even with good tasks, review SEO pull requests more strictly than ordinary UI changes, because mistakes are quiet. A wrong canonical or a stray noindex doesn't throw an error. It slowly removes pages from search. What I look for:
- The diff stays inside the stated scope.
- No changes to
robots.txt, redirects,middleware, or sitemap generation unless the task asked for them. - The verification output is pasted into the PR description.
- Structured data passes Google's Rich Results Test for one example URL.
After deploying, re-crawl the affected URLs and use URL Inspection in Search Console on one or two of them. Search engines take days to weeks to reprocess changes, so judge the fix by the HTML now and the rankings later.
Where an audit tool fits
Most of the work above is translation: turning "18 pages missing canonical" into one scoped, verifiable task. That's the part I got tired of doing by hand, and it's why I built Siteory. It audits a site for SEO, AI-search readiness, and security, and every finding comes with evidence and a fix prompt you can paste into Claude, Codex, Cursor, or OpenCode. Paid reports also export Markdown files like prioritized-fixes.md. It doesn't edit your code. You review the change, ship it, and re-scan.
Whether you use a tool or a spreadsheet, the rules are the same. One cause per task, a hard scope, acceptance criteria you can see in the HTML, and a command that proves it. Agents are fast. Give them something they can be checked against.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.