From 10cd46ee238c41dc1c0a5ccf8701a85efc45bf83 Mon Sep 17 00:00:00 2001 From: Karim shoair Date: Wed, 11 Feb 2026 03:55:23 +0200 Subject: [PATCH] docs: add the first page of the spiders section --- docs/spiders/architecture.md | 196 +++++++++++++++++++++++++++++++++++ zensical.toml | 6 ++ 2 files changed, 202 insertions(+) create mode 100644 docs/spiders/architecture.md diff --git a/docs/spiders/architecture.md b/docs/spiders/architecture.md new file mode 100644 index 0000000..70fba8c --- /dev/null +++ b/docs/spiders/architecture.md @@ -0,0 +1,196 @@ +## Introduction + +!!! success "Prerequisites" + + 1. You've completed or read the [Fetchers basics](../fetching/choosing.md) page to understand the different fetcher types and when to use each one. + 2. You've completed or read the [Main classes](../parsing/main_classes.md) page to understand the [Selector](../parsing/main_classes.md#selector) and [Response](../fetching/choosing.md#response-object) classes. + +Scrapling's spider system is a Scrapy-inspired async crawling framework designed for concurrent, multi-session crawls with built-in pause/resume support. It brings together Scrapling's parsing engine and fetchers into a unified crawling API while adding scheduling, concurrency control, and checkpointing. + +If you're familiar with Scrapy, you'll feel right at home. If not, don't worry โ€” the system is designed to be straightforward. + +## Architecture Overview + +The diagram below shows how data flows through the spider system when a crawl is running: + +```mermaid +graph TB + subgraph Spider["๐Ÿ•ท๏ธ Spider"] + direction TB + SR["start_requests()"] + Parse["parse() / callbacks"] + Hooks["Lifecycle hooks
on_start ยท on_close ยท on_error
is_blocked ยท on_scraped_item
"] + end + + subgraph Engine["โš™๏ธ Crawler Engine"] + direction TB + Loop["Crawl loop"] + Conc["Concurrency control
Global limiter ยท Per-domain limiter
Download delay
"] + BlockDet["Blocked request
detection & retry"] + end + + subgraph Scheduler["๐Ÿ“‹ Scheduler"] + PQ["Priority queue"] + Dedup["Fingerprint
deduplication"] + end + + subgraph SessionMgr["๐Ÿ”Œ Session Manager"] + direction TB + Router["Request router
Routes by session ID"] + subgraph Sessions["Sessions"] + direction LR + HTTP["FetcherSession
HTTP requests"] + Dynamic["AsyncDynamicSession
Browser automation"] + Stealth["AsyncStealthySession
Anti-bot bypass"] + end + end + + subgraph Output["๐Ÿ“ฆ Output"] + Items["ItemList
to_json() ยท to_jsonl()"] + Stream["Streaming
async for item in spider.stream()"] + Stats["CrawlStats"] + end + + Checkpoint["๐Ÿ’พ Checkpoint
Pause/Resume"] + + %% Data flow + SR -- "1. Initial
Requests" --> Scheduler + Scheduler -- "2. Next
Request" --> Engine + Engine -- "3. Fetch" --> SessionMgr + Router --> Sessions + SessionMgr -- "4. Response" --> Engine + Engine -- "5. Response" --> Parse + Parse -- "6a. New Requests" --> Scheduler + Parse -- "6b. Items (dict)" --> Output + Engine -. "Periodic save" .-> Checkpoint + Checkpoint -. "Resume" .-> Scheduler + + %% Styling + classDef spiderClass fill:#7c4dff,stroke:#333,color:#fff + classDef engineClass fill:#00897b,stroke:#333,color:#fff + classDef schedulerClass fill:#1565c0,stroke:#333,color:#fff + classDef sessionClass fill:#ef6c00,stroke:#333,color:#fff + classDef outputClass fill:#2e7d32,stroke:#333,color:#fff + classDef checkpointClass fill:#6d4c41,stroke:#333,color:#fff + + class Spider spiderClass + class Engine engineClass + class Scheduler schedulerClass + class SessionMgr sessionClass + class Output outputClass + class Checkpoint checkpointClass +``` + +## Data Flow + +Here's what happens step by step when you run a spider: + +### 1. Spider generates initial Requests + +The **Spider** produces the first batch of `Request` objects. By default, it creates one request for each URL in `start_urls`, but you can override `start_requests()` for custom logic (e.g., POST requests, authentication, different callbacks). + +### 2. Requests are scheduled + +The **Scheduler** receives requests and places them in a priority queue. Before queuing, each request gets a fingerprint (SHA1 hash of URL + method + body + session ID). Duplicate requests are silently dropped unless `dont_filter=True` is set on the request. Higher-priority requests are dequeued first. + +### 3. Engine fetches via Session Manager + +The **Crawler Engine** dequeues the next request, respecting concurrency limits (global and per-domain) and download delays. It then hands the request to the **Session Manager**, which routes it to the correct session based on the request's `sid` (session ID). + +A spider can have multiple sessions running simultaneously โ€” for example, a fast `FetcherSession` for simple pages and a `AsyncStealthySession` for protected pages. The Session Manager handles starting, routing, and closing all of them. + +### 4. Response comes back + +The session fetches the page and returns a `Response` object. The engine records statistics (response bytes, status codes, per-session request counts) and checks for blocked responses. + +If the response is blocked (determined by `is_blocked()`), the engine retries the request up to `max_blocked_retries` times, calling `retry_blocked_request()` so you can modify the request before retrying (e.g., switch proxies, add headers). + +### 5. Spider processes the Response + +The engine passes the `Response` to the spider's callback (default: `parse()`). The callback is an async generator that yields results. + +### 6. Results are routed + +The callback can yield three types of values: + +- **`dict`** โ€” Treated as a scraped item. Passed through `on_scraped_item()` for processing/filtering (return `None` to drop it), then added to the `ItemList` or streamed to the consumer. +- **`Request`** โ€” A follow-up request. Checked against `allowed_domains`, normalized, and sent to the Scheduler for queuing. +- **`None`** โ€” Silently ignored. + +The cycle repeats from step 2 until the scheduler is empty and no tasks are active, or the spider is paused. + +### Checkpointing + +If `crawldir` is set, the engine periodically saves a checkpoint (pending requests + seen URLs set) to disk. On graceful shutdown (Ctrl+C), a final checkpoint is saved. The next time the spider runs with the same `crawldir`, it resumes from where it left off โ€” skipping `start_requests()` and restoring the scheduler state. + + +## Components + +### Spider + +The central class you interact with. You subclass `Spider`, define your `start_urls` and `parse()` method, and optionally configure sessions and override lifecycle hooks. + +```python +from scrapling.spiders import Spider, Response, Request + +class MySpider(Spider): + name = "my_spider" + start_urls = ["https://example.com"] + + async def parse(self, response: Response): + for link in response.css("a::attr(href)").getall(): + yield response.follow(link, callback=self.parse_page) + + async def parse_page(self, response: Response): + yield {"title": response.css("h1::text").get("")} +``` + +### Crawler Engine + +The engine orchestrates the entire crawl. It manages the main loop, enforces concurrency limits using `anyio` task groups and `CapacityLimiter`, dispatches requests through the Session Manager, and processes results from callbacks. You don't interact with it directly โ€” the `Spider.start()` and `Spider.stream()` methods handle it for you. + +### Scheduler + +A priority queue with built-in URL deduplication. Requests are fingerprinted based on their URL, HTTP method, body, and session ID. The scheduler supports `snapshot()` and `restore()` for the checkpoint system, allowing the crawl state to be saved and resumed. + +### Session Manager + +Manages one or more named session instances. Each session is one of: + +- **`FetcherSession`** โ€” Async HTTP requests via `curl_cffi` (fastest, no JavaScript) +- **`AsyncDynamicSession`** โ€” Playwright browser automation (JavaScript support) +- **`AsyncStealthySession`** โ€” Stealthy browser with anti-bot bypass (Cloudflare, etc.) + +When a request comes in, the Session Manager routes it to the correct session based on the request's `sid` field. Sessions can be started eagerly (default) or lazily (started on first use). + +### Checkpoint System + +Saves the crawler's state (pending requests + seen URL fingerprints) to a pickle file on disk. Writes are atomic (temp file + rename) to prevent corruption. Checkpoints are saved periodically at a configurable interval and on graceful shutdown. On successful completion (not paused), checkpoint files are cleaned up automatically. + +### Output + +Scraped items are collected in an `ItemList` (a list subclass with `to_json()` and `to_jsonl()` export methods). Crawl statistics are tracked in a `CrawlStats` dataclass with per-domain bytes, per-session request counts, status code distribution, and timing information. + + +## Comparison with Scrapy + +If you're coming from Scrapy, here's how Scrapling's spider system maps: + +| Concept | Scrapy | Scrapling | +|---------|--------|-----------| +| Spider definition | `scrapy.Spider` subclass | `scrapling.spiders.Spider` subclass | +| Initial requests | `start_requests()` | `start_requests()` | +| Callbacks | `def parse(self, response)` | `async def parse(self, response)` | +| Following links | `response.follow(url)` | `response.follow(url)` | +| Item output | `yield dict` or `yield Item` | `yield dict` | +| Request scheduling | Scheduler + Dupefilter | Scheduler with built-in deduplication | +| Downloading | Downloader + Middlewares | Session Manager with multi-session support | +| Item processing | Item Pipelines | `on_scraped_item()` hook | +| Blocked detection | Custom middlewares | Built-in `is_blocked()` + `retry_blocked_request()` | +| Concurrency | `CONCURRENT_REQUESTS` setting | `concurrent_requests` class attribute | +| Domain filtering | `allowed_domains` | `allowed_domains` | +| Pause/Resume | `JOBDIR` setting | `crawldir` constructor argument | +| Export | Feed exports | `result.items.to_json()` / `to_jsonl()` | +| Running | `scrapy crawl spider_name` | `MySpider().start()` | +| Streaming | N/A | `async for item in spider.stream()` | +| Multi-session | N/A | Multiple fetcher sessions per spider | diff --git a/zensical.toml b/zensical.toml index 841a462..e0b99f1 100644 --- a/zensical.toml +++ b/zensical.toml @@ -30,6 +30,9 @@ nav = [ {"Dynamic websites" = "fetching/dynamic.md"}, {"Dynamic websites with hard protections" = "fetching/stealthy.md"} ]}, + {Spiders = [ + {"Architecture" = "spiders/architecture.md"}, + ]}, {"Command Line Interface" = [ {Overview = "cli/overview.md"}, {"Interactive shell" = "cli/interactive-shell.md"}, @@ -111,6 +114,9 @@ toggle.name = "Switch to light mode" #[project.markdown_extensions.mkautodoc] [project.markdown_extensions.pymdownx.details] [project.markdown_extensions.pymdownx.superfences] +custom_fences = [ + {name = "mermaid", class = "mermaid"} +] [project.markdown_extensions.pymdownx.inlinehilite] [project.markdown_extensions.pymdownx.snippets] [project.markdown_extensions.tables]