SEO Agent

Open-source SEO audit agent — real-browser checks, backlink scoring, and multi-engine AI citation tracking (Claude, ChatGPT, Gemini, Perplexity, Copilot) via SearchApi

Digital Marketing & SEOPythonMIT

Abstract

SEO Agent is an open-source Digital Marketing & SEO project. Open-source SEO audit agent — real-browser checks, backlink scoring, and multi-engine AI citation tracking (Claude, ChatGPT, Gemini, Perplexity, Copilot) via SearchApi. A local SEO co-pilot built with Python, Browser Use, and the Claude API. Visits real pages in a visible browser window, extracts SEO signals, checks for broken links, scores backlinks, surfaces GSC quick wins, maps internal link clusters, and writes structured reports — resumable if interrupted. It is built using Python. Key capabilities include: Visits each URL in a real Chromium browser — not a headless scraper; Extracts title, meta description, H1s, and canonical tag via Claude API; Checks for broken same-domain links asynchronously using httpx. The complete source code is publicly available on GitHub under the MIT License, making it a useful reference for students building a Digital Marketing & SEO mini project or final-year project.

1. Introduction

A local SEO co-pilot built with Python, Browser Use, and the Claude API. Visits real pages in a visible browser window, extracts SEO signals, checks for broken links, scores backlinks, surfaces GSC quick wins, maps internal link clusters, and writes structured reports — resumable if interrupted.

Ran it on my own sites. Found a title cannibalising its own homepage, a position 9.5 query with 0% CTR, two missing internal links, and an orphan page with no path to it.

2. Objective

Open-source SEO audit agent — real-browser checks, backlink scoring, and multi-engine AI citation tracking (Claude, ChatGPT, Gemini, Perplexity, Copilot) via SearchApi

This project demonstrates how Python can be applied to a real-world Digital Marketing & SEO problem.

3. Key Features / Modules

  • Visits each URL in a real Chromium browser — not a headless scraper
  • Extracts title, meta description, H1s, and canonical tag via Claude API
  • Checks for broken same-domain links asynchronously using httpx
  • Detects edge cases (404s, login walls, redirects) and pauses for human input
  • Writes results to report.json incrementally — safe to interrupt and resume
  • Generates a plain-English report-summary.txt on completion
  • qualify-backlinks — score a list of referring domains for niche relevance and traffic quality
  • gsc-insights — parse a Search Console export and find quick wins and cannibalisation
  • relevance-score — score candidate pages as internal link sources for a target URL
  • cluster-audit — map your full site into topic clusters, find orphans and missing hubs

4. Technology Stack

Python
  • Browser Use — real browser navigation via Playwright
  • Anthropic Claude API — structured SEO signal extraction (Haiku for modules, Sonnet for core)
  • SearchApi — structured Google SERP data for feature detection, plus ChatGPT, Gemini, Perplexity, and Bing Copilot answer-engine queries for LLM visibility
  • Python 3.11+, flat JSON state files, no database required

5. System Requirements

General requirements for this technology stack — check the README for exact versions.

  • Python 3.8 or later
  • pip / virtualenv for dependencies
  • VS Code, PyCharm or Jupyter Notebook
  • Git (to clone the repository)

6. Installation & Setup

git clone https://github.com/dannwaneri/seo-agent.git
cd seo-agent
git clone https://github.com/dannwaneri/seo-agent
cd seo-agent
pip install -r requirements.txt
playwright install chromium

Full setup instructions are in the project README.

7. Future Enhancements

Suggested extensions you can add to make this your own project.

  • Export reports to Google Sheets or PDF
  • Schedule weekly automated reports
  • Add competitor comparison

8. Viva / Review Questions

Common questions examiners ask for projects in this domain.

  1. Which marketing or SEO problem does this tool solve?
  2. Which data sources or APIs does it use (Search Console, Analytics, social platforms)?
  3. Which metrics or KPIs does it report and how are they calculated?
  4. How could the output help a business make decisions?
  5. How would you schedule it to run automatically?

9. Source Code & License

This project is developed by dannwaneri and published on GitHub under the MIT License. Please follow the license terms and credit the original author when you use or modify this code.

Want to build this as your internship project?

Work on a Digital Marketing & SEO project like this with mentor guidance, weekly reviews and an internship certificate from Training Trains, Erode — online or offline.

Apply for Digital Marketing & SEO Internship