Skip to main content
ScraperGuard - Scraping Observability & Drift Detection

ScraperGuard - Scraping Observability & Drift Detection

Production-grade observability layer for web scraping pipelines featuring deterministic root-cause attribution, DOM structural diffing, null-ratio drift detection, and composite health scoring. The reliability layer scrapers never knew they needed.

June 1, 2025 - Present
20 min read
Solo Creator & Maintainer
Developer Tools Β· Data Infrastructure
Python Library / CLI / REST API / Docker
Solo (I designed, built, and maintain everything)
Python FastAPI Pydantic v2 lxml SQLite Docker Click CLI Scrapy Playwright GitHub Actions MyPy REST API

Repository Stats

11
Stars
1
Forks
11
Watchers
3
Issues
Primary Language: Python

Overview

Web scrapers don’t crash β€” they silently drift. An HTTP 200 response looks successful, but behind the scenes, selectors break, fields go null, layouts restructure, and corrupted data flows downstream for days before anyone notices.

I built ScraperGuard to solve this exact problem: a production-grade observability and reliability layer that sits between your scraping pipeline and your data consumers, continuously monitoring extraction quality and pinpointing the exact root cause when things go wrong β€” all without touching a single line of your existing scraper code.

This isn’t another scraping framework. It’s the reliability layer that every scraping pipeline needs but never had.

8
Failure Types Classified
~5,800
Lines of Production Code
100%
Type Coverage (Strict MyPy)
0
Lines of Your Code Changed
Without ScraperGuard
Days of silent data corruption
From blind to observable
With ScraperGuard
Minutes to root-cause diagnosis
When a website redesigns its price element and 80% of prices go null, traditional scrapers ship corrupted data for days. ScraperGuard detects the drift within minutes and tells you exactly why: selector '.price-tag' broke due to DOM restructure, confidence 0.90.

The Problem I Solved

Why This Matters at Scale

In production scraping pipelines processing thousands of pages daily, the most dangerous failures aren’t the ones that throw errors β€” they’re the ones that return HTTP 200 with subtly wrong data. A selector that matched last week returns empty strings this week. A field that was always populated now has a 40% null rate. A CAPTCHA page gets parsed as product data. These silent failures corrupt downstream analytics, ML training data, pricing engines, and business decisions for days before anyone notices.

Traditional scraping tools focus exclusively on extraction and leave correctness completely unmeasured:

  • Silent Drift β€” Pages change without HTTP errors; corrupted data flows downstream undetected
  • Unclear Failures β€” When scraping breaks, diagnostics are minimal: β€œselector returned 0 matches”
  • Unmeasured Quality β€” No metrics on field completeness, null rate shifts, or structural stability
  • No Attribution β€” Is the failure a selector break? DOM restructure? CAPTCHA? Rate limiting? A/B test variant?

ScraperGuard’s unique value: Deterministic root-cause attribution β€” no ML black boxes, no guessing. Rule-based classification that identifies the exact failure type with evidence trails and recommended actions.


What I Owned (Solo Creator Scope)

I designed, architected, and built every component of ScraperGuard from scratch:

Architecture & System Design
Designed the full six-layer pipeline: Snapshot Capture β†’ Schema Validation β†’ DOM Diffing β†’ Failure Classification β†’ Health Scoring β†’ Alerting
Core Engine Development
Built the HTML normalizer, two-level fingerprinting system, tree-diff algorithm, and 8-type failure classifier with confidence scoring
Framework Integrations
Created non-invasive middleware for Scrapy and async context managers for Playwright β€” zero modifications to existing scraper code
API & CLI Tooling
Built a full FastAPI REST dashboard with 10+ endpoints and a Click CLI with run, validate, diff, report, and serve commands
Storage & Data Layer
Designed abstract storage interface with pluggable backends, implemented SQLite with WAL mode, foreign keys, and JSON columns
DevOps & Quality
Set up CI/CD with GitHub Actions, Docker containerization, 100% strict MyPy coverage, Ruff linting, and comprehensive test suite

Why I Chose This Stack

Every technology choice was driven by the specific requirements of a developer-facing observability tool: performance, type safety, zero-friction integration, and extensibility.

Python 3.11+ (Pure Python)

Scrapy and most scraping infrastructure runs on Python. Building ScraperGuard as a native Python library means zero-friction integration β€” pip install and you're done. No sidecar processes, no FFI, no language bridges.

Evaluated: Rust, Go, Node.js

Pydantic v2 for Schema Validation

Developers already know Pydantic. By inheriting from BaseSchema (which extends BaseModel), users get full type checking, custom validators, and automatic null tracking with zero boilerplate. Pydantic v2's Rust-powered core gives us validation speed that matches compiled languages.

Evaluated: JSON Schema, Marshmallow, attrs

lxml for DOM Processing

DOM structural diffing requires parsing thousands of HTML pages with complex structures. lxml's C-backed parser is 10-50x faster than html.parser or BeautifulSoup and handles malformed HTML gracefully β€” critical for real-world scraping where pages are rarely well-formed.

Evaluated: BeautifulSoup, html5lib, selectolax

SQLite with WAL Mode

ScraperGuard needs to store snapshots, validation results, and drift history without requiring users to set up external databases. SQLite with WAL mode gives us concurrent reads during writes, perfect for scraping pipelines that write snapshots while the API serves health reports.

Evaluated: PostgreSQL, MongoDB, JSONL files

Technical Architecture

The system employs a six-layer pipeline architecture where each layer has a single responsibility and is independently testable:


Technology Stack

Core Engine

Python 3.11+ β€” Modern Python with type hints, pattern matching, and performance improvements
Pydantic v2 β€” Rust-powered data validation for schema checking and field-level validators
lxml β€” High-performance C-backed HTML/XML parsing for DOM tree operations
SHA-256 Fingerprinting β€” Two-level hashing: structure-only and full-content for efficient change detection

Integrations

Scrapy Middleware β€” Non-invasive downloader middleware and item pipeline for automatic monitoring
Playwright Observer β€” Async context manager wrapping page navigation for browser-based scraping
FastAPI β€” REST API dashboard with 10+ endpoints for health reports, drift analysis, and selector tracking
Click CLI β€” Full command-line interface: run, validate, diff, report, and serve commands

Storage & Data

SQLite (WAL Mode) β€” Embedded database with concurrent reads/writes, foreign keys, and JSON columns
Abstract Storage Interface β€” Pluggable backend design supporting SQLite, PostgreSQL, and JSONL
YAML Configuration β€” 12-factor compatible config with environment variable interpolation

Alerting & Notifications

Slack Webhooks β€” Formatted block notifications for critical health events
HTTP Webhooks β€” Generic POST to custom monitoring endpoints
Severity Routing β€” Configurable thresholds: info, warning, critical with per-channel routing

Quality & DevOps

Strict MyPy β€” 100% type coverage with strict mode enabled across entire codebase
Ruff β€” Fast linting and formatting (E, F, I, N, W, UP rules)
GitHub Actions β€” CI/CD testing across Python 3.11, 3.12, and 3.13
Docker β€” Production-ready containerization with docker-compose for local development

Key Features & Capabilities

Two-Level Fingerprinting System

Most monitoring tools diff full HTML on every run β€” expensive and noisy. I designed a two-level fingerprinting system that dramatically reduces unnecessary computation:

  • Structure Fingerprint (SHA-256 of tag-nesting skeleton) β€” triggers expensive DOM diffing only when actual elements are added, removed, or moved
  • Content Fingerprint (SHA-256 of full normalized HTML) β€” fast β€œdid anything change?” check before deeper analysis
  • Smart noise removal: strips scripts, styles, tracking attributes, framework IDs, and inline styles before fingerprinting
  • Same fingerprint for <div>Price: $79</div> and <div>Price: $80</div> β€” text changes don’t trigger structural alerts

Deterministic Failure Classification

The classification engine identifies 8 distinct failure types with confidence scoring and evidence trails β€” no ML, fully debuggable:

╔═══════════════════════════════════════════════════════════════════════╗
β•‘  Health Score: 38/100 (CRITICAL)                                      β•‘
╠═══════════════════════════════════════════════════════════════════════╣
β•‘  Schema Compliance: 20.0% (6/30 items valid)                          β•‘
β•‘  Field Completeness: 56.0% (price weakest at 20%)                     β•‘
β•‘  Selector Stability: 50.0% (1/2 selectors broken)                     β•‘
β•‘  Structural Stability: 15.0% (4 DOM changes detected)                  β•‘
β•‘                                                                        β•‘
β•‘  FAILURES:                                                             β•‘
β•‘    [critical] SELECTOR_BREAK: '.price-tag' returns 0 matches           β•‘
β•‘               Confidence: 0.90                                         β•‘
β•‘               Action: Update CSS selector for price field              β•‘
β•‘                                                                        β•‘
β•‘  DRIFT ALERTS:                                                         β•‘
β•‘    [critical] price: Null ratio 3.3% β†’ 80.0% (+76.7%)                β•‘
β•‘    [warning]  rating: Null ratio 1.2% β†’ 18.5% (+17.3%)               β•‘
β•šβ•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•β•
from scraperguard.schema import BaseSchema, validators
 
class ProductSchema(BaseSchema):
    """Define what valid data looks like β€” ScraperGuard enforces it."""
    title: str                                      # Required string
    price: float = validators.range(min=0.01)       # Must be > $0.01
    availability: bool                               # Boolean field
    rating: float = validators.range(min=0, max=5)  # Between 0-5
    sku: str = validators.pattern(r"^[A-Z]{2}-\d{6}$")  # Regex pattern
  • SELECTOR_BREAK (0.90 confidence) β€” Selector returns 0 matches on an otherwise valid page
  • DOM_RESTRUCTURE (0.85) β€” 5+ high-severity changes or 3+ node removals detected
  • CAPTCHA_INJECTION (0.75–0.92) β€” Known CAPTCHA signatures (reCAPTCHA, hCaptcha, Cloudflare) detected in response
  • JS_CHALLENGE (0.80) β€” Minimal visible text with script-heavy response indicating browser rendering required
  • RATE_LIMIT (0.95) β€” HTTP 429 or rate-limit message patterns in response body
  • AB_VARIANT (0.75) β€” Partial selector failure with moderate DOM changes suggesting A/B test variant
  • PARTIAL_EXTRACTION (0.80) β€” Some items valid, others fail validation within the same batch
  • EMPTY_RESPONSE (0.95) β€” Response body under 100 bytes indicating server error or block

Non-Invasive Integration

The integration layer is designed to be completely transparent to existing scraper code β€” zero modifications required:

# settings.py β€” Just add two lines to your existing Scrapy project
DOWNLOADER_MIDDLEWARES = {
    "scraperguard.integrations.scrapy.ObserverMiddleware": 543,
}
ITEM_PIPELINES = {
    "scraperguard.integrations.scrapy.ValidationPipeline": 300,
}
 
# Configure what to monitor
SCRAPERGUARD_SCHEMA = "myproject.schemas.ProductSchema"
SCRAPERGUARD_SELECTORS = [".price", ".title", ".availability"]
 
# That's it. Your spider code stays exactly the same.
# ScraperGuard captures snapshots, validates items, and alerts
# on drift β€” all without touching your extraction logic.
from scraperguard.integrations.playwright import observe
 
# Wrap your existing Playwright code with observe()
async with observe(page, storage=storage, schema=ProductSchema) as observer:
    await page.goto("https://example.com/products")
 
    # Extract data exactly as you already do
    items = await extract_items(page)
 
    # Tell the observer what you extracted
    observer.set_items(items)
 
# On context exit: ScraperGuard automatically captures page.content(),
# runs the full pipeline (validate β†’ diff β†’ classify β†’ score),
# and returns a HealthReport. All failures caught internally β€”
# never crashes your browser automation.

Composite Health Scoring

A single 0–100 metric combining four independent quality signals, designed for dashboards, SLO tracking, and alerting thresholds:

  • Schema Compliance (30%) β€” Percentage of extracted items passing Pydantic validation
  • Field Completeness (30%) β€” Non-null ratios across all fields, penalizing the weakest field most heavily
  • Selector Stability (25%) β€” Percentage of tracked CSS selectors still finding matches
  • Structural Stability (15%) β€” DOM tree integrity based on severity-weighted change deductions
  • Three status levels: Healthy (β‰₯80), Degraded (50–79), Critical (<50)

Null Ratio Drift Detection

Statistical drift detection that catches subtle data quality regressions before they cascade downstream:

# ScraperGuard continuously compares current null ratios against
# the historical baseline (last N runs) for every field:
 
# Run 1-5 baseline: price null ratio = 3.3%
# Run 6 (current):  price null ratio = 80.0%
# Delta: +76.7% β†’ CRITICAL drift alert
 
# Severity thresholds (configurable):
#   > 50% delta β†’ CRITICAL (immediate action)
#   > 30% delta β†’ WARNING  (investigate soon)
#   > 15% delta β†’ INFO     (monitor closely)

Full CLI & REST API

# Run a full health check on a URL
scraperguard run https://example.com/products \
  --schema myschema.py \
  --selectors ".price,.title,.availability"
 
# Compare the last 2 DOM snapshots for a URL
scraperguard diff --url https://example.com --last 2
 
# Export a health report as JSON
scraperguard report --url https://example.com --format json
 
# Start the FastAPI dashboard
scraperguard serve --host 0.0.0.0 --port 8000
GET  /api/health                    Service health check
GET  /api/runs                      List recent scraper runs
GET  /api/runs/{run_id}             Run details with classifications
GET  /api/snapshots?url=...         Snapshots for a specific URL
GET  /api/validation/{url}          Validation result history
GET  /api/drift/{url}               Null ratio drift trends
GET  /api/report/{run_id}           Full health report with evidence
GET  /api/selectors/{url}           Selector stability across runs

Technical Challenges & Solutions

Challenge 1: Efficient DOM Diffing at Scale

The Challenge

Comparing raw HTML strings is both expensive and noisy β€” whitespace changes, dynamic attributes (data-reactid, style tags), and tracking scripts generate thousands of false positives. At scale, diffing every page on every run would be computationally prohibitive and produce unusable noise.

The Solution

I designed a multi-stage approach: first, normalize HTML by stripping noise elements (scripts, styles, comments), removing noisy attributes (event handlers, tracking IDs, framework attributes, inline styles), and normalizing whitespace. Then, generate structure-only fingerprints (SHA-256 of tag-nesting skeleton) that ignore text content entirely. DOM diffing only triggers when the structure fingerprint changes β€” text-only changes like price updates skip the expensive tree comparison entirely. The tree diff itself uses a greedy child-matching algorithm (O(nΒ²) per level) with tag + class similarity scoring, detecting 7 distinct change types with severity assignments.

# Two-level fingerprinting: skip expensive diffing when only text changed
structure_fp = fingerprint_structure(normalized_html)  # Tag skeleton only
content_fp = fingerprint_content(normalized_html)       # Full content
 
if structure_fp == previous_structure_fp:
    # DOM structure unchanged β€” skip tree diffing entirely
    # Only text/content changed (e.g., prices updated)
    structural_stability = 1.0
else:
    # Structure changed β€” run full tree diff
    changes = diff_trees(previous_dom, current_dom)
    # Returns: NODE_REMOVED, NODE_ADDED, TAG_CHANGED,
    #          ATTRIBUTES_CHANGED, CHILDREN_REORDERED, TEXT_CHANGED

Challenge 2: Root-Cause Attribution Without Machine Learning

The Challenge

When scraping fails, the typical diagnostic is minimal: β€œgot 0 results.” But the root cause could be a selector break, a site redesign, a CAPTCHA challenge, rate limiting, an A/B test variant, or a JavaScript-rendered page. ML-based classification would be a black box β€” teams need to understand and trust the diagnosis to act on it in production.

The Solution

I built a deterministic rule-based classifier with explicit detection signals for each of 8 failure types. Each classification includes a confidence score (0.0–1.0) based on signal strength, an evidence trail explaining what was detected, affected fields mapping, and a recommended remediation action. The classifier runs checks in priority order (empty response and rate limits first, then CAPTCHAs and JS challenges, then selector breaks and DOM restructures), returning only the highest-confidence classification per failure. Because it’s purely rule-based, every diagnosis is reproducible, debuggable, and explainable.

# Deterministic detection β€” every signal is explicit and traceable
captcha_signatures = [
    "g-recaptcha", "h-captcha", "captcha-container",
    "recaptcha/api.js", "hcaptcha.com", "challenge-platform",
    "cf-challenge", "cloudflare"
]
 
if any(sig in response_body for sig in captcha_signatures):
    return Classification(
        failure_type=FailureType.CAPTCHA_INJECTION,
        confidence=0.92,
        evidence=["CAPTCHA signature detected: 'g-recaptcha'"],
        recommended_action="Implement CAPTCHA solving or rotate proxy",
        severity="critical"
    )
# No ML. No guessing. Fully debuggable in production.

Challenge 3: Non-Invasive Integration with Existing Pipelines

The Challenge

Scraping teams have complex, battle-tested pipelines they’ve refined over years. Any observability tool that requires restructuring their code, modifying their spiders, or changing their data flow will face immediate rejection. The tool needs to be invisible to existing scraper logic while capturing everything needed for analysis.

The Solution

For Scrapy, I implemented a downloader middleware that intercepts responses without modifying them, storing metadata in request.meta as a side-channel. The item pipeline validates extracted items after the spider processes them β€” the spider never knows ScraperGuard exists. For Playwright, I built an async context manager that wraps page navigation and captures page.content() on exit. All internal failures are caught and logged β€” ScraperGuard never crashes the spider or browser automation, even if its own analysis fails. The result: two lines of config for Scrapy, one context manager for Playwright, zero changes to extraction logic.


Challenge 4: Meaningful Health Metrics from Noisy Signals

The Challenge

Individual quality signals (schema validation, null ratios, selector matches, DOM changes) each tell a partial story. A 95% schema compliance rate sounds healthy, but if the 5% failures are all in the price field, the data is commercially useless. A single DOM change could be trivial (text update) or catastrophic (layout restructure). Combining these into a single actionable metric requires careful weighting and normalization.

The Solution

I designed a 4-component weighted health score (0–100) where each component captures a different quality dimension: Schema Compliance (30%) measures type correctness, Field Completeness (30%) measures null ratios with penalty weighting for the weakest field, Selector Stability (25%) measures extraction reliability, and Structural Stability (15%) measures DOM integrity with severity-weighted deductions. The weights reflect real-world impact β€” a selector break (25% weight) is more actionable than a minor DOM change (15% weight), but data completeness (30% weight) is the strongest signal of downstream value. Three status levels (Healthy β‰₯80, Degraded 50–79, Critical <50) map directly to alerting thresholds.


Performance & Design Optimizations

Two-Level Fingerprinting

Skip expensive DOM diffing when only text content changes β€” structure fingerprints detect when tree comparison is actually needed

Greedy Child Matching

O(nΒ²) per tree level using tag + class similarity β€” fast enough for real-time monitoring without expensive sequence alignment algorithms

WAL Mode SQLite

Write-Ahead Logging enables concurrent reads during writes β€” scraping pipeline writes snapshots while API serves health reports simultaneously

HTML Noise Removal

Strip scripts, styles, tracking attributes, and framework IDs before fingerprinting β€” eliminates false positives from dynamic page elements

Batch Validation

Validate entire item batches in a single pass with aggregated null ratio computation β€” avoid per-item overhead

Pluggable Architecture

Abstract interfaces for storage, alerting, and classification β€” extend any layer without modifying core engine code


Results & Impact

8 Failure Types
Classified Automatically

From selector breaks to CAPTCHA injection β€” each with confidence scoring, evidence trails, and recommended actions

100%
Type Coverage

Strict MyPy across the entire codebase β€” every function, every parameter, every return type is verified at compile time

0 Lines
Of Your Code Changed

Non-invasive middleware and context managers integrate with existing Scrapy and Playwright projects without touching scraper logic

Open Source
MIT Licensed

Production-grade quality released as open source β€” comprehensive test suite, Docker support, and CI/CD across Python 3.11-3.13


Project Metrics

CategoryMetric
Codebase~5,800 lines of production code
~2,100 lines of test code
100% strict MyPy type coverage
CI/CD across Python 3.11, 3.12, 3.13
Core Engine8 deterministic failure classifications
7 DOM change type detections
4-component weighted health scoring
Two-level fingerprinting (structure + content)
IntegrationsScrapy downloader middleware + item pipeline
Playwright async context manager
FastAPI REST API with 10+ endpoints
Click CLI with 5 commands
AlertingSlack webhook notifications
Generic HTTP webhook support
Configurable severity thresholds
InfrastructureDocker + docker-compose containerization
SQLite with WAL mode + JSON columns
YAML config with env var interpolation
Pluggable storage backend architecture

Skills & Competencies Demonstrated

Systems Architecture

  • Designed a six-layer pipeline with strict separation of concerns
  • Created abstract interfaces for pluggable backends (storage, alerting, classification)
  • Built for extensibility without sacrificing simplicity
  • Each component independently testable and replaceable

Python Engineering

  • 100% strict MyPy coverage β€” every function signature, every return type
  • Pydantic v2 integration for zero-boilerplate schema validation
  • Async/await patterns for Playwright integration
  • Click CLI framework for developer-facing tooling
  • FastAPI for REST API with automatic OpenAPI documentation

Algorithm Design

  • Greedy child-matching tree diff algorithm for DOM structural comparison
  • Two-level SHA-256 fingerprinting for efficient change detection
  • Rule-based classification engine with confidence scoring and evidence trails
  • Weighted composite scoring with domain-specific weighting

Data Engineering

  • HTML normalization pipeline stripping 10+ noise categories
  • Null ratio drift detection with configurable statistical thresholds
  • Batch validation with aggregated field-level failure tracking
  • Time-series storage design for historical drift analysis

Developer Experience

  • Non-invasive integration: 2 lines of config for Scrapy, 1 context manager for Playwright
  • CLI with intuitive commands matching developer mental models
  • REST API for dashboard integration and programmatic access
  • YAML configuration with environment variable support and sensible defaults

DevOps & Quality

  • GitHub Actions CI/CD testing across 3 Python versions
  • Docker containerization with production-ready configuration
  • Comprehensive linting (Ruff) with strict rule enforcement
  • Test suite covering unit, integration, and example scenarios

Conclusion

ScraperGuard represents my approach to solving infrastructure problems: find a critical gap that everyone works around but nobody solves properly, and build a focused, well-engineered solution that integrates without friction.

Web scraping is a multi-billion dollar industry, yet the reliability tooling is nearly nonexistent. ScraperGuard brings the same observability mindset that transformed backend services (Datadog, Prometheus, PagerDuty) to the world of web scraping β€” where failures are silent, attribution is hard, and data quality degrades invisibly.

Designed solo. Engineered for production. Open source. The reliability layer every scraping pipeline needs.


Want to Work on Something Similar?

I'm available for freelance projects and full-time opportunities. Let's build something amazing together!