Skip to content
Boomspot
  • Home
Loading...
Boomspot

Daily tech news, software development coverage, Apple reporting, and the gear behind modern music making.

TwitterLinkedIn

Browse

  • Categories
  • Tags
  • Authors

Company

  • About
  • Contact

Legal

  • Privacy Policy
  • Terms of Service
  • Unsubscribe

© 2026 Boomspot. All rights reserved.

Built by Boomspot
Updated hourly

AI Content Disclosure: Articles on Boomspot are researched, written, and edited with the assistance of advanced AI systems. We combine software-assisted research with editorial oversight to deliver useful, accurate, and practical technical and music production content. Learn more about our editorial approach.

Browse by Category

Technology656Coding159Linux39SEO31Music Production24Studio Gear18Apple Rumors11

Popular Posts

Are Cracked VST Plugins Safe? A Producer's Reality Check

Are Cracked VST Plugins Safe? A Producer's Reality Check

6 min read
Shotcut vs Kdenlive: Best Free Linux Video Editor?

Shotcut vs Kdenlive: Best Free Linux Video Editor?

6 min read
Best Free DAW for Beginner Beatmakers: Full Comparison

Best Free DAW for Beginner Beatmakers: Full Comparison

6 min read
OLED vs QLED vs Mini-LED: Which TV Is Safest From Burn-In?

OLED vs QLED vs Mini-LED: Which TV Is Safest From Burn-In?

5 min read
How to Disable Firefox's Nova Redesign on Linux

How to Disable Firefox's Nova Redesign on Linux

5 min read

Recent Posts

Raspberry Pi 5 Alternatives After the $77.50 Price Hike

Raspberry Pi 5 Alternatives After the $77.50 Price Hike

Oct 1, 2026•6 min
PGO + LTO Explained: Do Compiler Tweaks Really Help?

PGO + LTO Explained: Do Compiler Tweaks Really Help?

Oct 1, 2026•7 min
GamesSoundCon 2026: Worth It for Composers?

GamesSoundCon 2026: Worth It for Composers?

Oct 1, 2026•6 min
Phobos vs Polaris vs Phoenix: Which BT Synth to Buy

Phobos vs Polaris vs Phoenix: Which BT Synth to Buy

Oct 1, 2026•7 min
Which Ethernet Cable Category Do You Actually Need?

Which Ethernet Cable Category Do You Actually Need?

Oct 1, 2026•6 min
  1. Home
  2. Coding
  3. Build a Python SEO Audit Script: Step-by-Step Guide
coding7 min read

Build a Python SEO Audit Script: Step-by-Step Guide

Stop clicking through pages one by one. This guide gives you a complete Python script that crawls your sitemap and outputs a prioritized CSV of every SEO issue it finds.

S

Staff

October 1, 2026

Reviewed byDorian

Build a Python SEO Audit Script: Step-by-Step Guide

If you run a multi-page site by yourself, you cannot afford to check every canonical tag, title length, and internal link by hand. A developer running a 100+ page coding platform found 42 pages pointing their canonical tags at the wrong domain, zero pages with structured data, and 50 pages with no internal links at all, all invisible until they wrote a script to check (dev.to). This guide gives you that script, fully built out, with the setup steps and verification checks the original post only sketched.

By the end, you will have a Python tool that reads your sitemap, checks each page for five common technical SEO problems, and writes a prioritized CSV you can hand to yourself or a teammate.

Step 1: Install the Right Dependencies

You need two packages: requests to fetch pages and beautifulsoup4 to parse HTML. Both are stable, well-documented, and require no API keys.

pip install requests==2.32.3 beautifulsoup4==4.12.3

Those are the versions confirmed compatible with Python 3.10 through 3.12 as of this writing. If you use a virtual environment (recommended), create one first with python -m venv venv and activate it before installing.

Step 2: Pull Every URL From Your Sitemap

Your sitemap.xml already lists every page you want indexed, so use it as the source of truth instead of manually listing URLs.

import requests
from xml.etree import ElementTree as ET

def get_urls_from_sitemap(sitemap_url):
    r = requests.get(sitemap_url, timeout=10)
    root = ET.fromstring(r.content)
    namespace = "{http://www.sitemaps.org/schemas/sitemap/0.9}"
    return [url.text for url in root.iter(f"{namespace}loc")]

This returns a plain list of strings. If your sitemap uses a sitemap index (a sitemap of sitemaps), you will need to loop through each child sitemap first, but most sites under a few hundred pages use a single flat file.

Step 3: Write the Core Audit Function

This is the unexecuted reference script. It checks the five issues most likely to silently kill your rankings: wrong canonicals, missing JSON-LD, oversized titles, thin internal linking, and thin content.

## audit.py — unexecuted reference script, verify locally before relying on results
import csv
import time
import requests
from bs4 import BeautifulSoup
from xml.etree import ElementTree as ET

YOUR_DOMAIN = "yoursite.com"  # replace with your real domain, no protocol

def get_urls_from_sitemap(sitemap_url):
    r = requests.get(sitemap_url, timeout=10)
    root = ET.fromstring(r.content)
    namespace = "{http://www.sitemaps.org/schemas/sitemap/0.9}"
    return [url.text for url in root.iter(f"{namespace}loc")]

def audit_page(url):
    issues = []
    try:
        r = requests.get(url, timeout=10)
    except requests.RequestException as e:
        return [{"url": url, "issue": f"Request failed: {e}", "severity": "critical"}] We cover related ground in [see also: how to fix flaky tests by managing application state](/how-to-fix-flaky-tests-by-managing-application-state).

    soup = BeautifulSoup(r.text, "html.parser")

    # Title length check
    title_tag = soup.title.string.strip() if soup.title and soup.title.string else ""
    if not title_tag:
        issues.append({"url": url, "issue": "Missing title tag", "severity": "critical"})
    elif len(title_tag) > 60:
        issues.append({"url": url, "issue": f"Title too long: {len(title_tag)} chars", "severity": "medium"})

    # Canonical check
    canonical = soup.find("link", attrs={"rel": "canonical"})
    if not canonical or not canonical.get("href"):
        issues.append({"url": url, "issue": "Missing canonical tag", "severity": "critical"})
    elif YOUR_DOMAIN not in canonical.get("href", ""):
        issues.append({"url": url, "issue": f"Wrong canonical domain: {canonical.get('href')}", "severity": "critical"})

    # JSON-LD structured data check
    if not soup.find("script", attrs={"type": "application/ld+json"}):
        issues.append({"url": url, "issue": "Missing JSON-LD structured data", "severity": "low"})

    # Internal link check
    internal_links = [a for a in soup.find_all("a", href=True)
                       if a["href"].startswith("/") or YOUR_DOMAIN in a["href"]]
    if len(internal_links) < 3:
        issues.append({"url": url, "issue": f"Few internal links: {len(internal_links)}", "severity": "high"})

    # Word count check
    text = soup.get_text(separator=" ", strip=True)
    word_count = len(text.split())
    if word_count < 300:
        issues.append({"url": url, "issue": f"Thin content: {word_count} words", "severity": "medium"})

    return issues if issues else [{"url": url, "issue": "No issues found", "severity": "none"}]

def run_audit(sitemap_url, output_file="seo_audit.csv"):
    urls = get_urls_from_sitemap(sitemap_url)
    all_issues = []

    for url in urls:
        page_issues = audit_page(url)
        all_issues.extend(page_issues)
        time.sleep(0.5)  # be polite to your own server

    severity_order = {"critical": 0, "high": 1, "medium": 2, "low": 3, "none": 4}
    all_issues.sort(key=lambda x: severity_order.get(x["severity"], 5))

    with open(output_file, "w", newline="", encoding="utf-8") as f:
        writer = csv.DictWriter(f, fieldnames=["url", "issue", "severity"])
        writer.writeheader()
        writer.writerows(all_issues)

    print(f"Audit complete. {len(urls)} pages checked. Results in {output_file}")

if __name__ == "__main__":
    run_audit("https://yoursite.com/sitemap.xml")

The time.sleep(0.5) line matters more than it looks. Hammering your own server with 200 rapid requests can trigger rate limiting on shared hosting, which will show up in your results as false "request failed" errors.

Step 4: Run It and Read the CSV

Save the script as audit.py, update YOUR_DOMAIN and the sitemap URL, then run:

python audit.py

The script writes seo_audit.csv, sorted so critical issues appear first. Here is what the output looks like for a small illustrative sample (not real crawl data, shown to demonstrate the CSV structure you should expect):

url,issue,severity
https://example.com/old-lesson,Wrong canonical domain: https://example.github.io/old-lesson,critical
https://example.com/new-page,Missing canonical tag,critical
https://example.com/snippet-3,Few internal links: 0,high
https://example.com/lesson-12,Thin content: 210 words,medium
https://example.com/how-to-loops,Title too long: 68 chars,medium
https://example.com/guide-1,Missing JSON-LD structured data,low
https://example.com/homepage,No issues found,none

Also read: our guide to slack code channels vs terminal agents: a decision guide

Open this in a spreadsheet, filter by severity, and work top-down. Canonical and missing-tag issues block indexing entirely; word count and JSON-LD issues affect ranking quality but won't stop Google from seeing the page.

Step 5: Verify the Script Actually Catches Problems

A script that always says "no issues found" is worse than no script. Before trusting the output across your whole site, test it against two pages you already understand.

  1. Point it at a page you know is clean: correct canonical, a title under 60 characters, JSON-LD present, several internal links, over 300 words. Confirm the CSV returns "No issues found" for that URL.
  2. Point it at a page you know is broken. Temporarily edit a test page's canonical tag to point at the wrong domain, or strip its internal links. Rerun the script and confirm it flags exactly that issue, nothing more, nothing less.

If the clean page throws false positives, check your YOUR_DOMAIN string. A mismatch between www.yoursite.com and yoursite.com in the canonical comparison is the most common cause of incorrect "wrong canonical" flags.

The One Failure Case You Need to Know About

This script uses requests, which fetches raw HTML only. It does not execute JavaScript. If your site is built with a client-side framework that injects title tags, canonical links, or content after the page loads in a browser, requests will see an empty shell and report false "missing title" or "thin content" errors on pages that are actually fine.

Check this yourself: run curl -s https://yoursite.com/some-page | grep title. If the title tag is missing from that raw output but visible in your browser, your site renders client-side and this script will misreport it. In that case, you need a headless browser tool like Playwright or Selenium instead of requests, which is a meaningfully bigger setup and outside the scope of this basic script.

Static-site generators, server-rendered frameworks, and most WordPress or Jekyll-style sites will not hit this problem. If you're unsure which category your site falls into, the curl test above answers it in ten seconds.

What to Expect Next

Once you have your first CSV, resist the urge to fix everything at once. Work through critical canonical and missing-tag issues first since those directly block indexing, then move to internal links, then title length and word count. Rerun the script weekly as you add new pages, since a fresh lesson or blog post with zero internal links will show up as an orphan the same way an old one would.

Tags

Software DevelopmentDeveloper ToolsCoding TutorialsCoding Best PracticesWeb Development

Keep reading

Unlocking ChatGPT Developer Mode: Full MCP Client Access
Coding•4 min read

Unlocking ChatGPT Developer Mode: Full MCP Client Access

Unlock the power of ChatGPT Developer Mode with full MCP client access. Discover how to enhance your coding projects and streamline development.

Sep 11, 2025

WebAssembly: Unleashing Native Speed in Web Browsers
Coding•4 min read

WebAssembly: Unleashing Native Speed in Web Browsers

WebAssembly is transforming web development with near-native performance, enabling more complex and efficient applications.

Sep 6, 2025

MCP Implementation at HubSpot: Elevating CRM with Context
Coding•4 min read

MCP Implementation at HubSpot: Elevating CRM with Context

Explore HubSpot's transformative MCP implementation for their CRM, detailing key strategies, challenges, and best practices for developers.

Sep 20, 2025

More stories for your next project

Get tech, coding, and music production updates in your inbox.

Unsubscribe anytime.