Dev.to AI 🤖 Ai 👁 0 📖 7 min read

Turning an SEO audit into tasks your AI coding agent can actually fix (Claude Code, Cursor, Codex)

A typical SEO audit export is a spreadsheet with a few hundred rows: "Missing canonical (18 pages)," "Duplicate title (11 pages)," "Orphan page," "Redirect chain." Pasting that into Claude Code, Cursor, or Codex with "fi

A typical SEO audit export is a spreadsheet with a few hundred rows: "Missing canonical (18 pages)," "Duplicate title (11 pages)," "Orphan page," "Redirect chain." Pasting that into Claude Code, Cursor, or Codex with "fix these" usually goes badly. The agent edits the wrong template, "fixes" a warning that didn't matter, or changes robots.txt in a way nobody reviews.

Coding agents are good at SEO fixes because most technical SEO problems are template problems, and a template is code. The trick is the translation step between the audit and the agent. This post is the process I use.

Step 1: Sort findings into three buckets

Not every SEO finding is a coding task. Before anything goes to an agent, I sort the list.

Delegate to the agent (deterministic, lives in code, easy to verify):

  • Missing, duplicate, or self-contradicting rel="canonical" tags
  • Duplicate or empty <title> and meta descriptions produced by a template
  • Leftover noindex (meta tag or X-Robots-Tag header) on production routes
  • Sitemap problems: wrong host, non-canonical URLs, URLs that 404 or redirect
  • Internal links pointing at redirects or 404s
  • Redirect chains (A→B→C) that should be a single hop
  • Missing or invalid JSON-LD (Organization, SoftwareApplication, BlogPosting, BreadcrumbList)
  • Missing alt attributes on content images, and heading-level order in components

Do it with the agent, but decide yourself (code changes that encode a business decision):

  • robots.txt rules, especially for AI crawlers
  • Which URL is canonical when two pages overlap
  • Redirect maps after a URL restructure
  • hreflang sets

Keep away from the agent (not code problems):

  • Content quality, search intent, and topic gaps
  • Backlinks and authority
  • Anything that's really a keyword-strategy decision

The first bucket usually covers most of what a technical audit flags. The second bucket is where an agent can do real damage if you let it decide on its own.

Step 2: Group by cause, not by URL

Audits report symptoms per URL. Agents fix causes. "Missing canonical on 18 pages" isn't 18 tasks. It's usually one layout or one generateMetadata-style function that never sets a canonical. Before writing tasks, look at the affected URLs and ask what they share: a route pattern (/blog/[slug]), a layout, a CMS content type.

One task per cause keeps diffs small and reviewable, and it stops the agent from patching 18 pages one by one, which you'd then have to maintain.

Step 3: Write each task so it can be verified

This template has worked well for me with all three agents:

## Task: Add self-referencing canonical to blog posts

**Finding:** 18 URLs under /blog/* have no <link rel="canonical">.
**Evidence:** `curl -s https://example.com/blog/hello-world | grep -i canonical` returns nothing.
**Likely cause:** app/blog/[slug]/page.tsx builds metadata without `alternates.canonical`.
**Change:** Set the canonical to the absolute production URL of the post
(https://example.com/blog/<slug>), built from the site's configured base URL.
**Scope:** Only files under app/blog/. Do not touch robots.txt, redirects, or other routes.
**Acceptance criteria:**
- Every /blog/<slug> page renders exactly one canonical tag.
- The canonical is absolute, uses https, the production host, and no query string.
- No other route's <head> output changes.
**Verify:** `npm run build`, then run scripts/check_seo.py against urls-blog.txt.

What each part does:

  • Evidence stops the agent from "fixing" something that isn't broken. If it can't reproduce the evidence, it should say so.
  • Likely cause is a hint, not an order. Agents are good at confirming or correcting it once they open the file.
  • Scope is the most important line. Without it, agents tidy up nearby code, and that's how a canonical fix turns into a 40-file diff.
  • Acceptance criteria are what you'll review against. Make them observable in rendered HTML, not in source code.
  • Verify gives the agent a way to check its own work before handing it back.

Step 4: Give the repo standing instructions

Rules that apply to every SEO task belong in the repository, not in each prompt. Codex and Cursor read an AGENTS.md file in the repo. Claude Code reads CLAUDE.md, and you can pull in a shared file from there with an @AGENTS.md import line. I keep the SEO rules in one place:

## SEO rules
- Production base URL: https://example.com (never localhost, preview, or staging hosts).
- Canonicals: absolute, https, production host, no query strings, self-referencing by default.
- Never add `noindex` or change robots.txt without an explicit instruction in the task.
- Titles come from the page's own data; never hardcode the site name as the full title.
- JSON-LD must only state facts visible on the page (no prices or ratings that aren't shown).
- After any <head> change, run: python3 scripts/check_seo.py urls.txt

That last rule matters most. It turns "looks right" into a command with an exit code.

Step 5: Verify against rendered HTML, not the diff

SEO bugs live in the HTML a crawler receives. A diff can look perfect while a parent layout overrides the canonical, or a client component injects a second <title>. So the check should fetch real pages. Here's a small standard-library Python script I give agents (and use myself). It takes a file of URLs, fails on missing or duplicate canonicals, canonicals pointing elsewhere, noindex, redirects, error status codes, and titles duplicated across pages:

#!/usr/bin/env python3
"""check_seo.py: verify basic SEO invariants for a list of URLs (one per line)."""
import sys
import urllib.error
import urllib.request
from urllib.parse import urlsplit, urlunsplit
from collections import defaultdict
from html.parser import HTMLParser


class HeadParser(HTMLParser):
    def __init__(self):
        super().__init__()
        self.titles, self.canonicals, self.robots = [], [], []
        self.in_title = self.in_body = False

    def handle_starttag(self, tag, attrs):
        a = {k: (v or "") for k, v in attrs}
        if tag == "body":
            self.in_body = True
        elif tag == "title" and not self.in_body:  # ignore <title> inside inline SVGs
            self.in_title = True
            self.titles.append("")
        elif tag == "link" and "canonical" in a.get("rel", "").lower().split():
            self.canonicals.append(a.get("href", ""))
        elif tag == "meta" and a.get("name", "").lower() == "robots":
            self.robots.append(a.get("content", ""))

    def handle_endtag(self, tag):
        if tag == "title":
            self.in_title = False

    def handle_data(self, data):
        if self.in_title:
            self.titles[-1] += data


def norm(u):
    """Treat https://a.com and https://a.com/ as the same URL."""
    parts = urlsplit(u)
    return urlunsplit(parts._replace(path=parts.path or "/"))


def check(url):
    req = urllib.request.Request(url, headers={"User-Agent": "seo-check/1.0"})
    try:
        with urllib.request.urlopen(req, timeout=15) as r:
            final, x_robots = r.geturl(), r.headers.get("X-Robots-Tag", "")
            html = r.read().decode("utf-8", "replace")
    except urllib.error.HTTPError as e:
        return None, [f"HTTP {e.code}"]
    p = HeadParser()
    p.feed(html)
    problems = []
    if norm(final) != norm(url):
        problems.append(f"redirects to {final}")
    if len(p.titles) != 1 or not p.titles[0].strip():
        problems.append(f"{len(p.titles)} <title> tags")
    if len(p.canonicals) != 1:
        problems.append(f"{len(p.canonicals)} canonical tags")
    elif norm(p.canonicals[0]) != norm(url):
        problems.append(f"canonical -> {p.canonicals[0]}")
    if any("noindex" in v.lower() for v in p.robots + [x_robots]):
        problems.append("noindex")
    title = p.titles[0].strip() if p.titles else None
    return title, problems


if __name__ == "__main__":
    urls = [line.strip() for line in open(sys.argv[1]) if line.strip()]
    by_title, failed = defaultdict(list), False
    for url in urls:
        title, problems = check(url)
        if title:
            by_title[title].append(url)
        print(("FAIL " if problems else "ok   ") + url + ("  " + "; ".join(problems) if problems else ""))
        failed |= bool(problems)
    for title, pages in by_title.items():
        if len(pages) > 1:
            print(f"FAIL duplicate title {title!r} on {len(pages)} URLs")
            failed = True
    sys.exit(1 if failed else 0)

Run it against a local production build (npm run build && npm start) with a URL list that covers one page per template, then again against production after deploy. List URLs in their canonical form. A trailing-slash mismatch counts as a failure on purpose, because it's a real inconsistency.

It's deliberately narrow. It doesn't judge title wording or content. It checks the invariants an agent is most likely to break.

Step 6: Review like it's a migration

Even with good tasks, review SEO pull requests more strictly than ordinary UI changes, because mistakes are quiet. A wrong canonical or a stray noindex doesn't throw an error. It slowly removes pages from search. What I look for:

  • The diff stays inside the stated scope.
  • No changes to robots.txt, redirects, middleware, or sitemap generation unless the task asked for them.
  • The verification output is pasted into the PR description.
  • Structured data passes Google's Rich Results Test for one example URL.

After deploying, re-crawl the affected URLs and use URL Inspection in Search Console on one or two of them. Search engines take days to weeks to reprocess changes, so judge the fix by the HTML now and the rankings later.

Where an audit tool fits

Most of the work above is translation: turning "18 pages missing canonical" into one scoped, verifiable task. That's the part I got tired of doing by hand, and it's why I built Siteory. It audits a site for SEO, AI-search readiness, and security, and every finding comes with evidence and a fix prompt you can paste into Claude, Codex, Cursor, or OpenCode. Paid reports also export Markdown files like prioritized-fixes.md. It doesn't edit your code. You review the change, ship it, and re-scan.

Whether you use a tool or a spreadsheet, the rules are the same. One cause per task, a hard scope, acceptance criteria you can see in the HTML, and a command that proves it. Agents are fast. Give them something they can be checked against.

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.