If you run a multi-page site by yourself, you cannot afford to check every canonical tag, title length, and internal link by hand. A developer running a 100+ page coding platform found 42 pages pointing their canonical tags at the wrong domain, zero pages with structured data, and 50 pages with no internal links at all, all invisible until they wrote a script to check (dev.to). This guide gives you that script, fully built out, with the setup steps and verification checks the original post only sketched.
By the end, you will have a Python tool that reads your sitemap, checks each page for five common technical SEO problems, and writes a prioritized CSV you can hand to yourself or a teammate.
Step 1: Install the Right Dependencies
You need two packages: requests to fetch pages and beautifulsoup4 to parse HTML. Both are stable, well-documented, and require no API keys.
pip install requests==2.32.3 beautifulsoup4==4.12.3
Those are the versions confirmed compatible with Python 3.10 through 3.12 as of this writing. If you use a virtual environment (recommended), create one first with python -m venv venv and activate it before installing.
Step 2: Pull Every URL From Your Sitemap
Your sitemap.xml already lists every page you want indexed, so use it as the source of truth instead of manually listing URLs.
import requests
from xml.etree import ElementTree as ET
def get_urls_from_sitemap(sitemap_url):
r = requests.get(sitemap_url, timeout=10)
root = ET.fromstring(r.content)
namespace = "{http://www.sitemaps.org/schemas/sitemap/0.9}"
return [url.text for url in root.iter(f"{namespace}loc")]
This returns a plain list of strings. If your sitemap uses a sitemap index (a sitemap of sitemaps), you will need to loop through each child sitemap first, but most sites under a few hundred pages use a single flat file.
Step 3: Write the Core Audit Function
This is the unexecuted reference script. It checks the five issues most likely to silently kill your rankings: wrong canonicals, missing JSON-LD, oversized titles, thin internal linking, and thin content.
## audit.py — unexecuted reference script, verify locally before relying on results
import csv
import time
import requests
from bs4 import BeautifulSoup
from xml.etree import ElementTree as ET
YOUR_DOMAIN = "yoursite.com" # replace with your real domain, no protocol
def get_urls_from_sitemap(sitemap_url):
r = requests.get(sitemap_url, timeout=10)
root = ET.fromstring(r.content)
namespace = "{http://www.sitemaps.org/schemas/sitemap/0.9}"
return [url.text for url in root.iter(f"{namespace}loc")]
def audit_page(url):
issues = []
try:
r = requests.get(url, timeout=10)
except requests.RequestException as e:
return [{"url": url, "issue": f"Request failed: {e}", "severity": "critical"}] We cover related ground in [see also: how to fix flaky tests by managing application state](/how-to-fix-flaky-tests-by-managing-application-state).
soup = BeautifulSoup(r.text, "html.parser")
# Title length check
title_tag = soup.title.string.strip() if soup.title and soup.title.string else ""
if not title_tag:
issues.append({"url": url, "issue": "Missing title tag", "severity": "critical"})
elif len(title_tag) > 60:
issues.append({"url": url, "issue": f"Title too long: {len(title_tag)} chars", "severity": "medium"})
# Canonical check
canonical = soup.find("link", attrs={"rel": "canonical"})
if not canonical or not canonical.get("href"):
issues.append({"url": url, "issue": "Missing canonical tag", "severity": "critical"})
elif YOUR_DOMAIN not in canonical.get("href", ""):
issues.append({"url": url, "issue": f"Wrong canonical domain: {canonical.get('href')}", "severity": "critical"})
# JSON-LD structured data check
if not soup.find("script", attrs={"type": "application/ld+json"}):
issues.append({"url": url, "issue": "Missing JSON-LD structured data", "severity": "low"})
# Internal link check
internal_links = [a for a in soup.find_all("a", href=True)
if a["href"].startswith("/") or YOUR_DOMAIN in a["href"]]
if len(internal_links) < 3:
issues.append({"url": url, "issue": f"Few internal links: {len(internal_links)}", "severity": "high"})
# Word count check
text = soup.get_text(separator=" ", strip=True)
word_count = len(text.split())
if word_count < 300:
issues.append({"url": url, "issue": f"Thin content: {word_count} words", "severity": "medium"})
return issues if issues else [{"url": url, "issue": "No issues found", "severity": "none"}]
def run_audit(sitemap_url, output_file="seo_audit.csv"):
urls = get_urls_from_sitemap(sitemap_url)
all_issues = []
for url in urls:
page_issues = audit_page(url)
all_issues.extend(page_issues)
time.sleep(0.5) # be polite to your own server
severity_order = {"critical": 0, "high": 1, "medium": 2, "low": 3, "none": 4}
all_issues.sort(key=lambda x: severity_order.get(x["severity"], 5))
with open(output_file, "w", newline="", encoding="utf-8") as f:
writer = csv.DictWriter(f, fieldnames=["url", "issue", "severity"])
writer.writeheader()
writer.writerows(all_issues)
print(f"Audit complete. {len(urls)} pages checked. Results in {output_file}")
if __name__ == "__main__":
run_audit("https://yoursite.com/sitemap.xml")
The time.sleep(0.5) line matters more than it looks. Hammering your own server with 200 rapid requests can trigger rate limiting on shared hosting, which will show up in your results as false "request failed" errors.
Step 4: Run It and Read the CSV
Save the script as audit.py, update YOUR_DOMAIN and the sitemap URL, then run:
python audit.py
The script writes seo_audit.csv, sorted so critical issues appear first. Here is what the output looks like for a small illustrative sample (not real crawl data, shown to demonstrate the CSV structure you should expect):
url,issue,severity
https://example.com/old-lesson,Wrong canonical domain: https://example.github.io/old-lesson,critical
https://example.com/new-page,Missing canonical tag,critical
https://example.com/snippet-3,Few internal links: 0,high
https://example.com/lesson-12,Thin content: 210 words,medium
https://example.com/how-to-loops,Title too long: 68 chars,medium
https://example.com/guide-1,Missing JSON-LD structured data,low
https://example.com/homepage,No issues found,none
Also read: our guide to slack code channels vs terminal agents: a decision guide
Open this in a spreadsheet, filter by severity, and work top-down. Canonical and missing-tag issues block indexing entirely; word count and JSON-LD issues affect ranking quality but won't stop Google from seeing the page.
Step 5: Verify the Script Actually Catches Problems
A script that always says "no issues found" is worse than no script. Before trusting the output across your whole site, test it against two pages you already understand.
- Point it at a page you know is clean: correct canonical, a title under 60 characters, JSON-LD present, several internal links, over 300 words. Confirm the CSV returns "No issues found" for that URL.
- Point it at a page you know is broken. Temporarily edit a test page's canonical tag to point at the wrong domain, or strip its internal links. Rerun the script and confirm it flags exactly that issue, nothing more, nothing less.
If the clean page throws false positives, check your YOUR_DOMAIN string. A mismatch between www.yoursite.com and yoursite.com in the canonical comparison is the most common cause of incorrect "wrong canonical" flags.
The One Failure Case You Need to Know About
This script uses requests, which fetches raw HTML only. It does not execute JavaScript. If your site is built with a client-side framework that injects title tags, canonical links, or content after the page loads in a browser, requests will see an empty shell and report false "missing title" or "thin content" errors on pages that are actually fine.
Check this yourself: run curl -s https://yoursite.com/some-page | grep title. If the title tag is missing from that raw output but visible in your browser, your site renders client-side and this script will misreport it. In that case, you need a headless browser tool like Playwright or Selenium instead of requests, which is a meaningfully bigger setup and outside the scope of this basic script.
Static-site generators, server-rendered frameworks, and most WordPress or Jekyll-style sites will not hit this problem. If you're unsure which category your site falls into, the curl test above answers it in ten seconds.
What to Expect Next
Once you have your first CSV, resist the urge to fix everything at once. Work through critical canonical and missing-tag issues first since those directly block indexing, then move to internal links, then title length and word count. Rerun the script weekly as you add new pages, since a fresh lesson or blog post with zero internal links will show up as an orphan the same way an old one would.



