From a4f31ae20553c8bd53d0be5126d8c4b14a811390 Mon Sep 17 00:00:00 2001 From: yetval Date: Thu, 19 Mar 2026 10:37:02 -0400 Subject: [PATCH] simple md fix --- .../Scrapling-Skill/references/fetching/dynamic.md | 2 +- .../Scrapling-Skill/references/spiders/architecture.md | 10 +++++----- .../references/spiders/getting-started.md | 2 +- .../Scrapling-Skill/references/spiders/sessions.md | 8 ++++---- 4 files changed, 11 insertions(+), 11 deletions(-) diff --git a/agent-skill/Scrapling-Skill/references/fetching/dynamic.md b/agent-skill/Scrapling-Skill/references/fetching/dynamic.md index a87fdf6..1a4c96d 100644 --- a/agent-skill/Scrapling-Skill/references/fetching/dynamic.md +++ b/agent-skill/Scrapling-Skill/references/fetching/dynamic.md @@ -44,7 +44,7 @@ Instead of launching a browser locally (Chromium/Google Chrome), you can connect **Notes:** * There was a `stealth` option here, but it was moved to the `StealthyFetcher` class, as explained on the next page, with additional features since version 0.3.13. -* This makes it less confusing for new users, easier to maintain, and provides other benefits, as explained on the [StealthyFetcher page](fetching/stealthy.md). +* This makes it less confusing for new users, easier to maintain, and provides other benefits, as explained on the [StealthyFetcher page](stealthy.md). ## Full list of arguments All arguments for `DynamicFetcher` and its session classes: diff --git a/agent-skill/Scrapling-Skill/references/spiders/architecture.md b/agent-skill/Scrapling-Skill/references/spiders/architecture.md index 9e7586f..de95d24 100644 --- a/agent-skill/Scrapling-Skill/references/spiders/architecture.md +++ b/agent-skill/Scrapling-Skill/references/spiders/architecture.md @@ -11,8 +11,8 @@ Here's what happens step by step when you run a spider: 1. The **Spider** produces the first batch of `Request` objects. By default, it creates one request for each URL in `start_urls`, but you can override `start_requests()` for custom logic. 2. The **Scheduler** receives requests and places them in a priority queue, and creates fingerprints for them. Higher-priority requests are dequeued first. 3. The **Crawler Engine** asks the **Scheduler** to dequeue the next request, respecting concurrency limits (global and per-domain) and download delays. Once the **Crawler Engine** receives the request, it passes it to the **Session Manager**, which routes it to the correct session based on the request's `sid` (session ID). -4. The **session** fetches the page and returns a [Response](fetching/choosing.md#response-object) object to the **Crawler Engine**. The engine records statistics and checks for blocked responses. If the response is blocked, the engine retries the request up to `max_blocked_retries` times. Of course, the blocking detection and the retry logic for blocked requests can be customized. -5. The **Crawler Engine** passes the [Response](fetching/choosing.md#response-object) to the request's callback. The callback either yields a dictionary, which gets treated as a scraped item, or a follow-up request, which gets sent to the scheduler for queuing. +4. The **session** fetches the page and returns a [Response](../fetching/choosing.md#response-object) object to the **Crawler Engine**. The engine records statistics and checks for blocked responses. If the response is blocked, the engine retries the request up to `max_blocked_retries` times. Of course, the blocking detection and the retry logic for blocked requests can be customized. +5. The **Crawler Engine** passes the [Response](../fetching/choosing.md#response-object) to the request's callback. The callback either yields a dictionary, which gets treated as a scraped item, or a follow-up request, which gets sent to the scheduler for queuing. 6. The cycle repeats from step 2 until the scheduler is empty and no tasks are active, or the spider is paused. 7. If `crawldir` is set while starting the spider, the **Crawler Engine** periodically saves a checkpoint (pending requests + seen URLs set) to disk. On graceful shutdown (Ctrl+C), a final checkpoint is saved. The next time the spider runs with the same `crawldir`, it resumes from where it left off — skipping `start_requests()` and restoring the scheduler state. @@ -50,9 +50,9 @@ A priority queue with built-in URL deduplication. Requests are fingerprinted bas Manages one or more named session instances. Each session is one of: -- [FetcherSession](fetching/static.md) -- [AsyncDynamicSession](fetching/dynamic.md) -- [AsyncStealthySession](fetching/stealthy.md) +- [FetcherSession](../fetching/static.md) +- [AsyncDynamicSession](../fetching/dynamic.md) +- [AsyncStealthySession](../fetching/stealthy.md) When a request comes in, the Session Manager routes it to the correct session based on the request's `sid` field. Sessions can be started with the spider start (default) or lazily (started on the first use). diff --git a/agent-skill/Scrapling-Skill/references/spiders/getting-started.md b/agent-skill/Scrapling-Skill/references/spiders/getting-started.md index 1446779..2a9e6c8 100644 --- a/agent-skill/Scrapling-Skill/references/spiders/getting-started.md +++ b/agent-skill/Scrapling-Skill/references/spiders/getting-started.md @@ -25,7 +25,7 @@ Every spider needs three things: 2. **`start_urls`** — A list of URLs to start crawling from. 3. **`parse()`** — An async generator method that processes each response and yields results. -The `parse()` method processes each response. You use the same selection methods you'd use with Scrapling's [Selector](parsing/main_classes.md#selector)/[Response](fetching/choosing.md#response-object), and `yield` dictionaries to output scraped items. +The `parse()` method processes each response. You use the same selection methods you'd use with Scrapling's [Selector](../parsing/main_classes.md#selector)/[Response](../fetching/choosing.md#response-object), and `yield` dictionaries to output scraped items. ## Running the Spider diff --git a/agent-skill/Scrapling-Skill/references/spiders/sessions.md b/agent-skill/Scrapling-Skill/references/spiders/sessions.md index 82a21ac..761cea6 100644 --- a/agent-skill/Scrapling-Skill/references/spiders/sessions.md +++ b/agent-skill/Scrapling-Skill/references/spiders/sessions.md @@ -6,14 +6,14 @@ A spider can use multiple fetcher sessions simultaneously — for example, a fas A session is a pre-configured fetcher instance that stays alive for the duration of the crawl. Instead of creating a new connection or browser for every request, the spider reuses sessions, which is faster and more resource-efficient. -By default, every spider creates a single [FetcherSession](fetching/static.md). You can add more sessions or swap the default by overriding the `configure_sessions()` method, but you have to use the async version of each session only, as the table shows below: +By default, every spider creates a single [FetcherSession](../fetching/static.md). You can add more sessions or swap the default by overriding the `configure_sessions()` method, but you have to use the async version of each session only, as the table shows below: | Session Type | Use Case | |-------------------------------------------------|------------------------------------------| -| [FetcherSession](fetching/static.md) | Fast HTTP requests, no JavaScript | -| [AsyncDynamicSession](fetching/dynamic.md) | Browser automation, JavaScript rendering | -| [AsyncStealthySession](fetching/stealthy.md) | Anti-bot bypass, Cloudflare, etc. | +| [FetcherSession](../fetching/static.md) | Fast HTTP requests, no JavaScript | +| [AsyncDynamicSession](../fetching/dynamic.md) | Browser automation, JavaScript rendering | +| [AsyncStealthySession](../fetching/stealthy.md) | Anti-bot bypass, Cloudflare, etc. | ## Configuring Sessions