ScraperGuard - Scraping Observability & Drift Detection
Production-grade observability layer for web scraping pipelines featuring deterministic root-cause attribution, DOM structural diffing, null-ratio drift detection, and composite health scoring. The reliability layer scrapers never knew they needed.
Repository Stats
Overview
Web scrapers donβt crash β they silently drift. An HTTP 200 response looks successful, but behind the scenes, selectors break, fields go null, layouts restructure, and corrupted data flows downstream for days before anyone notices.
I built ScraperGuard to solve this exact problem: a production-grade observability and reliability layer that sits between your scraping pipeline and your data consumers, continuously monitoring extraction quality and pinpointing the exact root cause when things go wrong β all without touching a single line of your existing scraper code.
This isnβt another scraping framework. Itβs the reliability layer that every scraping pipeline needs but never had.
The Problem I Solved
In production scraping pipelines processing thousands of pages daily, the most dangerous failures arenβt the ones that throw errors β theyβre the ones that return HTTP 200 with subtly wrong data. A selector that matched last week returns empty strings this week. A field that was always populated now has a 40% null rate. A CAPTCHA page gets parsed as product data. These silent failures corrupt downstream analytics, ML training data, pricing engines, and business decisions for days before anyone notices.
Traditional scraping tools focus exclusively on extraction and leave correctness completely unmeasured:
- Silent Drift β Pages change without HTTP errors; corrupted data flows downstream undetected
- Unclear Failures β When scraping breaks, diagnostics are minimal: βselector returned 0 matchesβ
- Unmeasured Quality β No metrics on field completeness, null rate shifts, or structural stability
- No Attribution β Is the failure a selector break? DOM restructure? CAPTCHA? Rate limiting? A/B test variant?
ScraperGuardβs unique value: Deterministic root-cause attribution β no ML black boxes, no guessing. Rule-based classification that identifies the exact failure type with evidence trails and recommended actions.
What I Owned (Solo Creator Scope)
I designed, architected, and built every component of ScraperGuard from scratch:
Why I Chose This Stack
Every technology choice was driven by the specific requirements of a developer-facing observability tool: performance, type safety, zero-friction integration, and extensibility.
Python 3.11+ (Pure Python)
Scrapy and most scraping infrastructure runs on Python. Building ScraperGuard as a native Python library means zero-friction integration β pip install and you're done. No sidecar processes, no FFI, no language bridges.
Pydantic v2 for Schema Validation
Developers already know Pydantic. By inheriting from BaseSchema (which extends BaseModel), users get full type checking, custom validators, and automatic null tracking with zero boilerplate. Pydantic v2's Rust-powered core gives us validation speed that matches compiled languages.
lxml for DOM Processing
DOM structural diffing requires parsing thousands of HTML pages with complex structures. lxml's C-backed parser is 10-50x faster than html.parser or BeautifulSoup and handles malformed HTML gracefully β critical for real-world scraping where pages are rarely well-formed.
SQLite with WAL Mode
ScraperGuard needs to store snapshots, validation results, and drift history without requiring users to set up external databases. SQLite with WAL mode gives us concurrent reads during writes, perfect for scraping pipelines that write snapshots while the API serves health reports.
Technical Architecture
The system employs a six-layer pipeline architecture where each layer has a single responsibility and is independently testable:
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β OBSERVER LAYER β
β βββββββββββββββββββ βββββββββββββββββββ βββββββββββββββββββ β
β β Scrapy Middlewareβ β Playwright β β Standalone β β
β β (Downloader + β β Page Observer β β Python API β β
β β Pipeline) β β (Async Context)β β (Direct Call) β β
β βββββββββββββββββββ βββββββββββββββββββ βββββββββββββββββββ β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β
Non-Invasive Capture (Zero Code Changes)
β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β SNAPSHOT CAPTURE ENGINE β
β βββββββββββββββββββ βββββββββββββββββββ βββββββββββββββββββ β
β β HTML Normalizer β β Two-Level β β Snapshot β β
β β (Noise Removal) β β Fingerprinting β β Persistence β β
β βββββββββββββββββββ βββββββββββββββββββ βββββββββββββββββββ β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β
Normalized HTML + Fingerprints
β
ββββββββββββββββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββββββββββ
β SCHEMA VALIDATION β DOM STRUCTURAL DIFF β
β βββββββββββββββββββ β βββββββββββββββββββ β
β β Pydantic v2 β β β Tree-Level β β
β β Type Checking β β β Comparison β β
β βββββββββββββββββββ€ β βββββββββββββββββββ€ β
β β Null Ratio β β β 7 Change Type β β
β β Drift Detection β β β Detection β β
β βββββββββββββββββββ€ β βββββββββββββββββββ€ β
β β Field-Level β β β Selector β β
β β Validators β β β Affinity Map β β
β βββββββββββββββββββ β βββββββββββββββββββ β
ββββββββββββββββββββββββββββββββββββ΄βββββββββββββββββββββββββββββββββββββββ
β
Validation Results + DOM Changes
β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β FAILURE CLASSIFICATION ENGINE β
β ββββββββββββββ ββββββββββββββ ββββββββββββββ ββββββββββββββ β
β β Selector β β DOM β β CAPTCHA β β Rate β β
β β Break β β Restructureβ β Injection β β Limit β β
β ββββββββββββββ€ ββββββββββββββ€ ββββββββββββββ€ ββββββββββββββ€ β
β β JS β β A/B β β Partial β β Empty β β
β β Challenge β β Variant β β Extraction β β Response β β
β ββββββββββββββ ββββββββββββββ ββββββββββββββ ββββββββββββββ β
β Confidence Scoring + Evidence Trails β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β
Classifications + Evidence
β
ββββββββββββββββββββββββββββββββββββ¬βββββββββββββββββββββββββββββββββββββββ
β HEALTH SCORING β ALERTING SYSTEM β
β βββββββββββββββββββ β βββββββββββββββββββ β
β β 4-Component β β β Slack Webhook β β
β β Weighted Score β β βββββββββββββββββββ€ β
β β (0-100 Scale) β β β HTTP Webhook β β
β βββββββββββββββββββ β βββββββββββββββββββ€ β
β Schema: 30% β Fields: 30% β β Email (SMTP) β β
β Selectors: 25% β DOM: 15% β βββββββββββββββββββ β
ββββββββββββββββββββββββββββββββββββ΄βββββββββββββββββββββββββββββββββββββββ
Technology Stack
Core Engine
Integrations
Storage & Data
Alerting & Notifications
Quality & DevOps
Key Features & Capabilities
Two-Level Fingerprinting System
Most monitoring tools diff full HTML on every run β expensive and noisy. I designed a two-level fingerprinting system that dramatically reduces unnecessary computation:
- Structure Fingerprint (SHA-256 of tag-nesting skeleton) β triggers expensive DOM diffing only when actual elements are added, removed, or moved
- Content Fingerprint (SHA-256 of full normalized HTML) β fast βdid anything change?β check before deeper analysis
- Smart noise removal: strips scripts, styles, tracking attributes, framework IDs, and inline styles before fingerprinting
- Same fingerprint for
<div>Price: $79</div>and<div>Price: $80</div>β text changes donβt trigger structural alerts
Deterministic Failure Classification
The classification engine identifies 8 distinct failure types with confidence scoring and evidence trails β no ML, fully debuggable:
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Health Score: 38/100 (CRITICAL) β
β ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ£
β Schema Compliance: 20.0% (6/30 items valid) β
β Field Completeness: 56.0% (price weakest at 20%) β
β Selector Stability: 50.0% (1/2 selectors broken) β
β Structural Stability: 15.0% (4 DOM changes detected) β
β β
β FAILURES: β
β [critical] SELECTOR_BREAK: '.price-tag' returns 0 matches β
β Confidence: 0.90 β
β Action: Update CSS selector for price field β
β β
β DRIFT ALERTS: β
β [critical] price: Null ratio 3.3% β 80.0% (+76.7%) β
β [warning] rating: Null ratio 1.2% β 18.5% (+17.3%) β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββfrom scraperguard.schema import BaseSchema, validators
class ProductSchema(BaseSchema):
"""Define what valid data looks like β ScraperGuard enforces it."""
title: str # Required string
price: float = validators.range(min=0.01) # Must be > $0.01
availability: bool # Boolean field
rating: float = validators.range(min=0, max=5) # Between 0-5
sku: str = validators.pattern(r"^[A-Z]{2}-\d{6}$") # Regex pattern- SELECTOR_BREAK (0.90 confidence) β Selector returns 0 matches on an otherwise valid page
- DOM_RESTRUCTURE (0.85) β 5+ high-severity changes or 3+ node removals detected
- CAPTCHA_INJECTION (0.75β0.92) β Known CAPTCHA signatures (reCAPTCHA, hCaptcha, Cloudflare) detected in response
- JS_CHALLENGE (0.80) β Minimal visible text with script-heavy response indicating browser rendering required
- RATE_LIMIT (0.95) β HTTP 429 or rate-limit message patterns in response body
- AB_VARIANT (0.75) β Partial selector failure with moderate DOM changes suggesting A/B test variant
- PARTIAL_EXTRACTION (0.80) β Some items valid, others fail validation within the same batch
- EMPTY_RESPONSE (0.95) β Response body under 100 bytes indicating server error or block
Non-Invasive Integration
The integration layer is designed to be completely transparent to existing scraper code β zero modifications required:
# settings.py β Just add two lines to your existing Scrapy project
DOWNLOADER_MIDDLEWARES = {
"scraperguard.integrations.scrapy.ObserverMiddleware": 543,
}
ITEM_PIPELINES = {
"scraperguard.integrations.scrapy.ValidationPipeline": 300,
}
# Configure what to monitor
SCRAPERGUARD_SCHEMA = "myproject.schemas.ProductSchema"
SCRAPERGUARD_SELECTORS = [".price", ".title", ".availability"]
# That's it. Your spider code stays exactly the same.
# ScraperGuard captures snapshots, validates items, and alerts
# on drift β all without touching your extraction logic.from scraperguard.integrations.playwright import observe
# Wrap your existing Playwright code with observe()
async with observe(page, storage=storage, schema=ProductSchema) as observer:
await page.goto("https://example.com/products")
# Extract data exactly as you already do
items = await extract_items(page)
# Tell the observer what you extracted
observer.set_items(items)
# On context exit: ScraperGuard automatically captures page.content(),
# runs the full pipeline (validate β diff β classify β score),
# and returns a HealthReport. All failures caught internally β
# never crashes your browser automation.Composite Health Scoring
A single 0β100 metric combining four independent quality signals, designed for dashboards, SLO tracking, and alerting thresholds:
- Schema Compliance (30%) β Percentage of extracted items passing Pydantic validation
- Field Completeness (30%) β Non-null ratios across all fields, penalizing the weakest field most heavily
- Selector Stability (25%) β Percentage of tracked CSS selectors still finding matches
- Structural Stability (15%) β DOM tree integrity based on severity-weighted change deductions
- Three status levels: Healthy (β₯80), Degraded (50β79), Critical (<50)
Null Ratio Drift Detection
Statistical drift detection that catches subtle data quality regressions before they cascade downstream:
# ScraperGuard continuously compares current null ratios against
# the historical baseline (last N runs) for every field:
# Run 1-5 baseline: price null ratio = 3.3%
# Run 6 (current): price null ratio = 80.0%
# Delta: +76.7% β CRITICAL drift alert
# Severity thresholds (configurable):
# > 50% delta β CRITICAL (immediate action)
# > 30% delta β WARNING (investigate soon)
# > 15% delta β INFO (monitor closely)Full CLI & REST API
# Run a full health check on a URL
scraperguard run https://example.com/products \
--schema myschema.py \
--selectors ".price,.title,.availability"
# Compare the last 2 DOM snapshots for a URL
scraperguard diff --url https://example.com --last 2
# Export a health report as JSON
scraperguard report --url https://example.com --format json
# Start the FastAPI dashboard
scraperguard serve --host 0.0.0.0 --port 8000GET /api/health Service health check
GET /api/runs List recent scraper runs
GET /api/runs/{run_id} Run details with classifications
GET /api/snapshots?url=... Snapshots for a specific URL
GET /api/validation/{url} Validation result history
GET /api/drift/{url} Null ratio drift trends
GET /api/report/{run_id} Full health report with evidence
GET /api/selectors/{url} Selector stability across runsTechnical Challenges & Solutions
Challenge 1: Efficient DOM Diffing at Scale
The Challenge
Comparing raw HTML strings is both expensive and noisy β whitespace changes, dynamic attributes (data-reactid, style tags), and tracking scripts generate thousands of false positives. At scale, diffing every page on every run would be computationally prohibitive and produce unusable noise.
The Solution
I designed a multi-stage approach: first, normalize HTML by stripping noise elements (scripts, styles, comments), removing noisy attributes (event handlers, tracking IDs, framework attributes, inline styles), and normalizing whitespace. Then, generate structure-only fingerprints (SHA-256 of tag-nesting skeleton) that ignore text content entirely. DOM diffing only triggers when the structure fingerprint changes β text-only changes like price updates skip the expensive tree comparison entirely. The tree diff itself uses a greedy child-matching algorithm (O(nΒ²) per level) with tag + class similarity scoring, detecting 7 distinct change types with severity assignments.
# Two-level fingerprinting: skip expensive diffing when only text changed
structure_fp = fingerprint_structure(normalized_html) # Tag skeleton only
content_fp = fingerprint_content(normalized_html) # Full content
if structure_fp == previous_structure_fp:
# DOM structure unchanged β skip tree diffing entirely
# Only text/content changed (e.g., prices updated)
structural_stability = 1.0
else:
# Structure changed β run full tree diff
changes = diff_trees(previous_dom, current_dom)
# Returns: NODE_REMOVED, NODE_ADDED, TAG_CHANGED,
# ATTRIBUTES_CHANGED, CHILDREN_REORDERED, TEXT_CHANGEDChallenge 2: Root-Cause Attribution Without Machine Learning
The Challenge
When scraping fails, the typical diagnostic is minimal: βgot 0 results.β But the root cause could be a selector break, a site redesign, a CAPTCHA challenge, rate limiting, an A/B test variant, or a JavaScript-rendered page. ML-based classification would be a black box β teams need to understand and trust the diagnosis to act on it in production.
The Solution
I built a deterministic rule-based classifier with explicit detection signals for each of 8 failure types. Each classification includes a confidence score (0.0β1.0) based on signal strength, an evidence trail explaining what was detected, affected fields mapping, and a recommended remediation action. The classifier runs checks in priority order (empty response and rate limits first, then CAPTCHAs and JS challenges, then selector breaks and DOM restructures), returning only the highest-confidence classification per failure. Because itβs purely rule-based, every diagnosis is reproducible, debuggable, and explainable.
# Deterministic detection β every signal is explicit and traceable
captcha_signatures = [
"g-recaptcha", "h-captcha", "captcha-container",
"recaptcha/api.js", "hcaptcha.com", "challenge-platform",
"cf-challenge", "cloudflare"
]
if any(sig in response_body for sig in captcha_signatures):
return Classification(
failure_type=FailureType.CAPTCHA_INJECTION,
confidence=0.92,
evidence=["CAPTCHA signature detected: 'g-recaptcha'"],
recommended_action="Implement CAPTCHA solving or rotate proxy",
severity="critical"
)
# No ML. No guessing. Fully debuggable in production.Challenge 3: Non-Invasive Integration with Existing Pipelines
The Challenge
Scraping teams have complex, battle-tested pipelines theyβve refined over years. Any observability tool that requires restructuring their code, modifying their spiders, or changing their data flow will face immediate rejection. The tool needs to be invisible to existing scraper logic while capturing everything needed for analysis.
The Solution
For Scrapy, I implemented a downloader middleware that intercepts responses without modifying them, storing metadata in request.meta as a side-channel. The item pipeline validates extracted items after the spider processes them β the spider never knows ScraperGuard exists. For Playwright, I built an async context manager that wraps page navigation and captures page.content() on exit. All internal failures are caught and logged β ScraperGuard never crashes the spider or browser automation, even if its own analysis fails. The result: two lines of config for Scrapy, one context manager for Playwright, zero changes to extraction logic.
Challenge 4: Meaningful Health Metrics from Noisy Signals
The Challenge
Individual quality signals (schema validation, null ratios, selector matches, DOM changes) each tell a partial story. A 95% schema compliance rate sounds healthy, but if the 5% failures are all in the price field, the data is commercially useless. A single DOM change could be trivial (text update) or catastrophic (layout restructure). Combining these into a single actionable metric requires careful weighting and normalization.
The Solution
I designed a 4-component weighted health score (0β100) where each component captures a different quality dimension: Schema Compliance (30%) measures type correctness, Field Completeness (30%) measures null ratios with penalty weighting for the weakest field, Selector Stability (25%) measures extraction reliability, and Structural Stability (15%) measures DOM integrity with severity-weighted deductions. The weights reflect real-world impact β a selector break (25% weight) is more actionable than a minor DOM change (15% weight), but data completeness (30% weight) is the strongest signal of downstream value. Three status levels (Healthy β₯80, Degraded 50β79, Critical <50) map directly to alerting thresholds.
Performance & Design Optimizations
Two-Level Fingerprinting
Skip expensive DOM diffing when only text content changes β structure fingerprints detect when tree comparison is actually needed
Greedy Child Matching
O(nΒ²) per tree level using tag + class similarity β fast enough for real-time monitoring without expensive sequence alignment algorithms
WAL Mode SQLite
Write-Ahead Logging enables concurrent reads during writes β scraping pipeline writes snapshots while API serves health reports simultaneously
HTML Noise Removal
Strip scripts, styles, tracking attributes, and framework IDs before fingerprinting β eliminates false positives from dynamic page elements
Batch Validation
Validate entire item batches in a single pass with aggregated null ratio computation β avoid per-item overhead
Pluggable Architecture
Abstract interfaces for storage, alerting, and classification β extend any layer without modifying core engine code
Results & Impact
From selector breaks to CAPTCHA injection β each with confidence scoring, evidence trails, and recommended actions
Strict MyPy across the entire codebase β every function, every parameter, every return type is verified at compile time
Non-invasive middleware and context managers integrate with existing Scrapy and Playwright projects without touching scraper logic
Production-grade quality released as open source β comprehensive test suite, Docker support, and CI/CD across Python 3.11-3.13
Project Metrics
| Category | Metric |
|---|---|
| Codebase | ~5,800 lines of production code |
| ~2,100 lines of test code | |
| 100% strict MyPy type coverage | |
| CI/CD across Python 3.11, 3.12, 3.13 | |
| Core Engine | 8 deterministic failure classifications |
| 7 DOM change type detections | |
| 4-component weighted health scoring | |
| Two-level fingerprinting (structure + content) | |
| Integrations | Scrapy downloader middleware + item pipeline |
| Playwright async context manager | |
| FastAPI REST API with 10+ endpoints | |
| Click CLI with 5 commands | |
| Alerting | Slack webhook notifications |
| Generic HTTP webhook support | |
| Configurable severity thresholds | |
| Infrastructure | Docker + docker-compose containerization |
| SQLite with WAL mode + JSON columns | |
| YAML config with env var interpolation | |
| Pluggable storage backend architecture |
Skills & Competencies Demonstrated
Systems Architecture
- Designed a six-layer pipeline with strict separation of concerns
- Created abstract interfaces for pluggable backends (storage, alerting, classification)
- Built for extensibility without sacrificing simplicity
- Each component independently testable and replaceable
Python Engineering
- 100% strict MyPy coverage β every function signature, every return type
- Pydantic v2 integration for zero-boilerplate schema validation
- Async/await patterns for Playwright integration
- Click CLI framework for developer-facing tooling
- FastAPI for REST API with automatic OpenAPI documentation
Algorithm Design
- Greedy child-matching tree diff algorithm for DOM structural comparison
- Two-level SHA-256 fingerprinting for efficient change detection
- Rule-based classification engine with confidence scoring and evidence trails
- Weighted composite scoring with domain-specific weighting
Data Engineering
- HTML normalization pipeline stripping 10+ noise categories
- Null ratio drift detection with configurable statistical thresholds
- Batch validation with aggregated field-level failure tracking
- Time-series storage design for historical drift analysis
Developer Experience
- Non-invasive integration: 2 lines of config for Scrapy, 1 context manager for Playwright
- CLI with intuitive commands matching developer mental models
- REST API for dashboard integration and programmatic access
- YAML configuration with environment variable support and sensible defaults
DevOps & Quality
- GitHub Actions CI/CD testing across 3 Python versions
- Docker containerization with production-ready configuration
- Comprehensive linting (Ruff) with strict rule enforcement
- Test suite covering unit, integration, and example scenarios
Conclusion
ScraperGuard represents my approach to solving infrastructure problems: find a critical gap that everyone works around but nobody solves properly, and build a focused, well-engineered solution that integrates without friction.
Web scraping is a multi-billion dollar industry, yet the reliability tooling is nearly nonexistent. ScraperGuard brings the same observability mindset that transformed backend services (Datadog, Prometheus, PagerDuty) to the world of web scraping β where failures are silent, attribution is hard, and data quality degrades invisibly.
Designed solo. Engineered for production. Open source. The reliability layer every scraping pipeline needs.
Want to Work on Something Similar?
I'm available for freelance projects and full-time opportunities. Let's build something amazing together!