SEO Tools

On-page SEO analyzer and site auditor. Crawls websites to surface metadata, content, and technical SEO issues, with AI/GenAI content readiness checks.

Digital Marketing & SEOTypeScriptMIT

Abstract

SEO Tools is an open-source Digital Marketing & SEO project. On-page SEO analyzer and site auditor. Crawls websites to surface metadata, content, and technical SEO issues, with AI/GenAI content readiness checks. For each page with a flagged title or meta description, generates a rewritten suggestion using an LLM, grounded in the page's own heading structure and content excerpt. Provider-agnostic via an OpenAI-compatible client — point it at OpenAI or a local model (Ollama) through env vars: LLM_PROVIDER (openai | ollama), LLM_API_KEY, LLM_MODEL, LLM_BASE_URL. It is built using TypeScript. Key capabilities include: Sliding Time Windows: More accurate than fixed time periods; Multiple Concurrent Rules: Apply several limits simultaneously; Persistent Storage: Request history survives crawler restarts. The complete source code is publicly available on GitHub under the MIT License, making it a useful reference for students building a Digital Marketing & SEO mini project or final-year project.

1. Introduction

For each page with a flagged title or meta description, generates a rewritten suggestion using an LLM, grounded in the page's own heading structure and content excerpt. Provider-agnostic via an OpenAI-compatible client — point it at OpenAI or a local model (Ollama) through env vars: LLM_PROVIDER (openai | ollama), LLM_API_KEY, LLM_MODEL, LLM_BASE_URL.

An on-page SEO analyzer that crawls a site and turns the result into actionable audits: SEO issue reports (titles, meta descriptions, canonicals, structured data, broken links), a comprehensive Markdown audit, optional LLM-powered title/meta-description fix suggestions, and an MCP server that exposes it all to AI agents. The Crawlee + Playwright crawler underneath extracts Google-supported SEO tags, structured data (JSON-LD/microdata), and AI-indexing metadata; the reporting layer is where the value is.

The typical workflow is crawl once → run reports — see Reports & Analysis.

2. Objective

On-page SEO analyzer and site auditor. Crawls websites to surface metadata, content, and technical SEO issues, with AI/GenAI content readiness checks.

This project demonstrates how TypeScript can be applied to a real-world Digital Marketing & SEO problem.

3. Key Features / Modules

  • Sliding Time Windows: More accurate than fixed time periods
  • Multiple Concurrent Rules: Apply several limits simultaneously
  • Persistent Storage: Request history survives crawler restarts
  • Smart Distribution: Even request spacing to avoid bursts
  • Real-time Monitoring: Status updates every 10 requests
  • Automatic Delays: Built-in waiting when limits are reached
  • Basic data: Title, URL, timestamp
  • Response data: HTTP status, headers
  • Links: Internal/external link analysis
  • SEO tags: Google meta tags

4. Technology Stack

TypeScript
  • Monitor HTTP response codes and redirects
  • Analyze response headers for performance insights
  • Track site structure and internal linking
  • Identify crawl errors and accessibility issues

5. System Requirements

General requirements for this technology stack — check the README for exact versions.

  • Node.js (LTS) and npm / yarn / pnpm
  • VS Code or any code editor
  • Git (to clone the repository)

6. Installation & Setup

git clone https://github.com/siva01c/seo-tools.git
cd seo-tools
# 1. Create your .env from the template (required first step), then edit values as needed
cp .env.example .env

# 2. Build the crawler image
docker compose build app

# 3. Crawl a site
docker compose run --rm app npm run crawl -- https://example.com --headless=true

# Crawl with exclusions and rate limiting
docker compose run --rm app npm run crawl -- https://example.com \
  --exclude-domains "api.example.com,cdn.example.com" \
  --rate-limit=conservative
# Crawl a blog for SEO analysis (with visible browser)
docker compose run --rm app npm run crawl -- https://myblog.com --headless=false

# Analyze e-commerce site structure (excluding API and CDN)
docker compose run --rm app npm run crawl -- https://mystore.com --exclude-domains "api.mystore.com,cdn.mystore.com"

# Debug crawling with visible browser
docker compose run --rm app npm run crawl -- https://example.com --headless=false

# Stealth crawling with invisible browser
docker compose run --rm app npm run crawl -- https://example.com --headless=true

# Test local development site
docker compose run --rm app npm run crawl -- http://localhost:3000 --headless=false

# Single page analysis with visible browser
docker compose run --rm app npm run crawl -- https://example.com --single --headless=false

# Visible browser with a domain exclusion
docker compose run --rm app npm run crawl -- https://www.example.com --headless=false --exclude-domains "accounts.example.com"

Full setup instructions are in the project README.

7. Future Enhancements

Suggested extensions you can add to make this your own project.

  • Export reports to Google Sheets or PDF
  • Schedule weekly automated reports
  • Add competitor comparison

8. Viva / Review Questions

Common questions examiners ask for projects in this domain.

  1. Which marketing or SEO problem does this tool solve?
  2. Which data sources or APIs does it use (Search Console, Analytics, social platforms)?
  3. Which metrics or KPIs does it report and how are they calculated?
  4. How could the output help a business make decisions?
  5. How would you schedule it to run automatically?

9. Source Code & License

This project is developed by siva01c and published on GitHub under the MIT License. Please follow the license terms and credit the original author when you use or modify this code.

Want to build this as your internship project?

Work on a Digital Marketing & SEO project like this with mentor guidance, weekly reviews and an internship certificate from Training Trains, Erode — online or offline.

Apply for Digital Marketing & SEO Internship