From 9af644ef84808466390544613c84cd01a3dbcd41 Mon Sep 17 00:00:00 2001 From: Karim shoair Date: Wed, 22 Apr 2026 16:52:59 +0200 Subject: [PATCH] docs: improving the code copy-paste experience and use less tokens for the agent skill --- .../references/fetching/choosing.md | 40 ++--- .../references/fetching/dynamic.md | 2 +- .../references/fetching/static.md | 146 ++++++++--------- .../references/fetching/stealthy.md | 2 +- .../references/parsing/adaptive.md | 41 +++-- .../references/parsing/main_classes.md | 24 +-- .../references/parsing/selection.md | 4 +- docs/cli/interactive-shell.md | 119 +++++++------- docs/development/scrapling_custom_types.md | 11 +- docs/fetching/choosing.md | 40 ++--- docs/fetching/dynamic.md | 2 +- docs/fetching/static.md | 148 +++++++++--------- docs/fetching/stealthy.md | 2 +- docs/overview.md | 58 ++++--- docs/parsing/adaptive.md | 41 +++-- docs/parsing/main_classes.md | 24 +-- docs/parsing/selection.md | 4 +- 17 files changed, 345 insertions(+), 363 deletions(-) diff --git a/agent-skill/Scrapling-Skill/references/fetching/choosing.md b/agent-skill/Scrapling-Skill/references/fetching/choosing.md index 10ec7e8..1ec0ef8 100644 --- a/agent-skill/Scrapling-Skill/references/fetching/choosing.md +++ b/agent-skill/Scrapling-Skill/references/fetching/choosing.md @@ -27,23 +27,23 @@ The following table compares them and can be quickly used for guidance. ## Parser configuration in all fetchers All fetchers share the same import method, as you will see in the upcoming pages ```python ->>> from scrapling.fetchers import Fetcher, AsyncFetcher, StealthyFetcher, DynamicFetcher +from scrapling.fetchers import Fetcher, AsyncFetcher, StealthyFetcher, DynamicFetcher ``` Then you use it right away without initializing like this, and it will use the default parser settings: ```python ->>> page = StealthyFetcher.fetch('https://example.com') +page = StealthyFetcher.fetch('https://example.com') ``` If you want to configure the parser ([Selector class](parsing/main_classes.md#selector)) that will be used on the response before returning it for you, then do this first: ```python ->>> from scrapling.fetchers import Fetcher ->>> Fetcher.configure(adaptive=True, keep_comments=False, keep_cdata=False) # and the rest +from scrapling.fetchers import Fetcher +Fetcher.configure(adaptive=True, keep_comments=False, keep_cdata=False) # and the rest ``` or ```python ->>> from scrapling.fetchers import Fetcher ->>> Fetcher.adaptive=True ->>> Fetcher.keep_comments=False ->>> Fetcher.keep_cdata=False # and the rest +from scrapling.fetchers import Fetcher +Fetcher.adaptive=True +Fetcher.keep_comments=False +Fetcher.keep_cdata=False # and the rest ``` Then, continue your code as usual. @@ -59,19 +59,19 @@ If your use case requires a different configuration for each request/fetch, you ## Response Object The `Response` object is the same as the [Selector](parsing/main_classes.md#selector) class, but it has additional details about the response, like response headers, status, cookies, etc., as shown below: ```python ->>> from scrapling.fetchers import Fetcher ->>> page = Fetcher.get('https://example.com') +from scrapling.fetchers import Fetcher +page = Fetcher.get('https://example.com') ->>> page.status # HTTP status code ->>> page.reason # Status message ->>> page.cookies # Response cookies as a dictionary ->>> page.headers # Response headers ->>> page.request_headers # Request headers ->>> page.history # Response history of redirections, if any ->>> page.body # Raw response body as bytes ->>> page.encoding # Response encoding ->>> page.meta # Response metadata dictionary (e.g., proxy used). Mainly helpful with the spiders system. ->>> page.captured_xhr # List of captured XHR/fetch responses (when capture_xhr is enabled on a browser session) +page.status # HTTP status code +page.reason # Status message +page.cookies # Response cookies as a dictionary +page.headers # Response headers +page.request_headers # Request headers +page.history # Response history of redirections, if any +page.body # Raw response body as bytes +page.encoding # Response encoding +page.meta # Response metadata dictionary (e.g., proxy used). Mainly helpful with the spiders system. +page.captured_xhr # List of captured XHR/fetch responses (when capture_xhr is enabled on a browser session) ``` All fetchers return the `Response` object. diff --git a/agent-skill/Scrapling-Skill/references/fetching/dynamic.md b/agent-skill/Scrapling-Skill/references/fetching/dynamic.md index a831d7c..5f5b0ce 100644 --- a/agent-skill/Scrapling-Skill/references/fetching/dynamic.md +++ b/agent-skill/Scrapling-Skill/references/fetching/dynamic.md @@ -8,7 +8,7 @@ As we will explain later, to automate the page, you need some knowledge of [Play You have one primary way to import this Fetcher, which is the same for all fetchers. ```python ->>> from scrapling.fetchers import DynamicFetcher +from scrapling.fetchers import DynamicFetcher ``` Check out how to configure the parsing options [here](choosing.md#parser-configuration-in-all-fetchers) diff --git a/agent-skill/Scrapling-Skill/references/fetching/static.md b/agent-skill/Scrapling-Skill/references/fetching/static.md index 4115483..e24f6ce 100644 --- a/agent-skill/Scrapling-Skill/references/fetching/static.md +++ b/agent-skill/Scrapling-Skill/references/fetching/static.md @@ -6,7 +6,7 @@ The `Fetcher` class provides rapid and lightweight HTTP requests using the high- Import the Fetcher (same import pattern for all fetchers): ```python ->>> from scrapling.fetchers import Fetcher +from scrapling.fetchers import Fetcher ``` Check out how to configure the parsing options [here](choosing.md#parser-configuration-in-all-fetchers) @@ -47,41 +47,41 @@ Examples are the best way to explain this: > Hence: `OPTIONS` and `HEAD` methods are not supported. #### GET ```python ->>> from scrapling.fetchers import Fetcher ->>> # Basic GET ->>> page = Fetcher.get('https://example.com') ->>> page = Fetcher.get('https://scrapling.requestcatcher.com/get', stealthy_headers=True) ->>> page = Fetcher.get('https://scrapling.requestcatcher.com/get', proxy='http://username:password@localhost:8030') ->>> # With parameters ->>> page = Fetcher.get('https://example.com/search', params={'q': 'query'}) ->>> ->>> # With headers ->>> page = Fetcher.get('https://example.com', headers={'User-Agent': 'Custom/1.0'}) ->>> # Basic HTTP authentication ->>> page = Fetcher.get("https://example.com", auth=("my_user", "password123")) ->>> # Browser impersonation ->>> page = Fetcher.get('https://example.com', impersonate='chrome') ->>> # HTTP/3 support ->>> page = Fetcher.get('https://example.com', http3=True) +from scrapling.fetchers import Fetcher +# Basic GET +page = Fetcher.get('https://example.com') +page = Fetcher.get('https://scrapling.requestcatcher.com/get', stealthy_headers=True) +page = Fetcher.get('https://scrapling.requestcatcher.com/get', proxy='http://username:password@localhost:8030') +# With parameters +page = Fetcher.get('https://example.com/search', params={'q': 'query'}) + +# With headers +page = Fetcher.get('https://example.com', headers={'User-Agent': 'Custom/1.0'}) +# Basic HTTP authentication +page = Fetcher.get("https://example.com", auth=("my_user", "password123")) +# Browser impersonation +page = Fetcher.get('https://example.com', impersonate='chrome') +# HTTP/3 support +page = Fetcher.get('https://example.com', http3=True) ``` And for asynchronous requests, it's a small adjustment ```python ->>> from scrapling.fetchers import AsyncFetcher ->>> # Basic GET ->>> page = await AsyncFetcher.get('https://example.com') ->>> page = await AsyncFetcher.get('https://scrapling.requestcatcher.com/get', stealthy_headers=True) ->>> page = await AsyncFetcher.get('https://scrapling.requestcatcher.com/get', proxy='http://username:password@localhost:8030') ->>> # With parameters ->>> page = await AsyncFetcher.get('https://example.com/search', params={'q': 'query'}) ->>> ->>> # With headers ->>> page = await AsyncFetcher.get('https://example.com', headers={'User-Agent': 'Custom/1.0'}) ->>> # Basic HTTP authentication ->>> page = await AsyncFetcher.get("https://example.com", auth=("my_user", "password123")) ->>> # Browser impersonation ->>> page = await AsyncFetcher.get('https://example.com', impersonate='chrome110') ->>> # HTTP/3 support ->>> page = await AsyncFetcher.get('https://example.com', http3=True) +from scrapling.fetchers import AsyncFetcher +# Basic GET +page = await AsyncFetcher.get('https://example.com') +page = await AsyncFetcher.get('https://scrapling.requestcatcher.com/get', stealthy_headers=True) +page = await AsyncFetcher.get('https://scrapling.requestcatcher.com/get', proxy='http://username:password@localhost:8030') +# With parameters +page = await AsyncFetcher.get('https://example.com/search', params={'q': 'query'}) + +# With headers +page = await AsyncFetcher.get('https://example.com', headers={'User-Agent': 'Custom/1.0'}) +# Basic HTTP authentication +page = await AsyncFetcher.get("https://example.com", auth=("my_user", "password123")) +# Browser impersonation +page = await AsyncFetcher.get('https://example.com', impersonate='chrome110') +# HTTP/3 support +page = await AsyncFetcher.get('https://example.com', http3=True) ``` The `page` object in all cases is a [Response](choosing.md#response-object) object, which is a [Selector](parsing/main_classes.md#selector), so you can use it directly ```python @@ -102,62 +102,62 @@ The `page` object in all cases is a [Response](choosing.md#response-object) obje ``` #### POST ```python ->>> from scrapling.fetchers import Fetcher ->>> # Basic POST ->>> page = Fetcher.post('https://scrapling.requestcatcher.com/post', data={'key': 'value'}, params={'q': 'query'}) ->>> page = Fetcher.post('https://scrapling.requestcatcher.com/post', data={'key': 'value'}, stealthy_headers=True) ->>> page = Fetcher.post('https://scrapling.requestcatcher.com/post', data={'key': 'value'}, proxy='http://username:password@localhost:8030', impersonate="chrome") ->>> # Another example of form-encoded data ->>> page = Fetcher.post('https://example.com/submit', data={'username': 'user', 'password': 'pass'}, http3=True) ->>> # JSON data ->>> page = Fetcher.post('https://example.com/api', json={'key': 'value'}) +from scrapling.fetchers import Fetcher +# Basic POST +page = Fetcher.post('https://scrapling.requestcatcher.com/post', data={'key': 'value'}, params={'q': 'query'}) +page = Fetcher.post('https://scrapling.requestcatcher.com/post', data={'key': 'value'}, stealthy_headers=True) +page = Fetcher.post('https://scrapling.requestcatcher.com/post', data={'key': 'value'}, proxy='http://username:password@localhost:8030', impersonate="chrome") +# Another example of form-encoded data +page = Fetcher.post('https://example.com/submit', data={'username': 'user', 'password': 'pass'}, http3=True) +# JSON data +page = Fetcher.post('https://example.com/api', json={'key': 'value'}) ``` And for asynchronous requests, it's a small adjustment ```python ->>> from scrapling.fetchers import AsyncFetcher ->>> # Basic POST ->>> page = await AsyncFetcher.post('https://scrapling.requestcatcher.com/post', data={'key': 'value'}) ->>> page = await AsyncFetcher.post('https://scrapling.requestcatcher.com/post', data={'key': 'value'}, stealthy_headers=True) ->>> page = await AsyncFetcher.post('https://scrapling.requestcatcher.com/post', data={'key': 'value'}, proxy='http://username:password@localhost:8030', impersonate="chrome") ->>> # Another example of form-encoded data ->>> page = await AsyncFetcher.post('https://example.com/submit', data={'username': 'user', 'password': 'pass'}, http3=True) ->>> # JSON data ->>> page = await AsyncFetcher.post('https://example.com/api', json={'key': 'value'}) +from scrapling.fetchers import AsyncFetcher +# Basic POST +page = await AsyncFetcher.post('https://scrapling.requestcatcher.com/post', data={'key': 'value'}) +page = await AsyncFetcher.post('https://scrapling.requestcatcher.com/post', data={'key': 'value'}, stealthy_headers=True) +page = await AsyncFetcher.post('https://scrapling.requestcatcher.com/post', data={'key': 'value'}, proxy='http://username:password@localhost:8030', impersonate="chrome") +# Another example of form-encoded data +page = await AsyncFetcher.post('https://example.com/submit', data={'username': 'user', 'password': 'pass'}, http3=True) +# JSON data +page = await AsyncFetcher.post('https://example.com/api', json={'key': 'value'}) ``` #### PUT ```python ->>> from scrapling.fetchers import Fetcher ->>> # Basic PUT ->>> page = Fetcher.put('https://example.com/update', data={'status': 'updated'}) ->>> page = Fetcher.put('https://example.com/update', data={'status': 'updated'}, stealthy_headers=True, impersonate="chrome") ->>> page = Fetcher.put('https://example.com/update', data={'status': 'updated'}, proxy='http://username:password@localhost:8030') ->>> # Another example of form-encoded data ->>> page = Fetcher.put("https://scrapling.requestcatcher.com/put", data={'key': ['value1', 'value2']}) +from scrapling.fetchers import Fetcher +# Basic PUT +page = Fetcher.put('https://example.com/update', data={'status': 'updated'}) +page = Fetcher.put('https://example.com/update', data={'status': 'updated'}, stealthy_headers=True, impersonate="chrome") +page = Fetcher.put('https://example.com/update', data={'status': 'updated'}, proxy='http://username:password@localhost:8030') +# Another example of form-encoded data +page = Fetcher.put("https://scrapling.requestcatcher.com/put", data={'key': ['value1', 'value2']}) ``` And for asynchronous requests, it's a small adjustment ```python ->>> from scrapling.fetchers import AsyncFetcher ->>> # Basic PUT ->>> page = await AsyncFetcher.put('https://example.com/update', data={'status': 'updated'}) ->>> page = await AsyncFetcher.put('https://example.com/update', data={'status': 'updated'}, stealthy_headers=True, impersonate="chrome") ->>> page = await AsyncFetcher.put('https://example.com/update', data={'status': 'updated'}, proxy='http://username:password@localhost:8030') ->>> # Another example of form-encoded data ->>> page = await AsyncFetcher.put("https://scrapling.requestcatcher.com/put", data={'key': ['value1', 'value2']}) +from scrapling.fetchers import AsyncFetcher +# Basic PUT +page = await AsyncFetcher.put('https://example.com/update', data={'status': 'updated'}) +page = await AsyncFetcher.put('https://example.com/update', data={'status': 'updated'}, stealthy_headers=True, impersonate="chrome") +page = await AsyncFetcher.put('https://example.com/update', data={'status': 'updated'}, proxy='http://username:password@localhost:8030') +# Another example of form-encoded data +page = await AsyncFetcher.put("https://scrapling.requestcatcher.com/put", data={'key': ['value1', 'value2']}) ``` #### DELETE ```python ->>> from scrapling.fetchers import Fetcher ->>> page = Fetcher.delete('https://example.com/resource/123') ->>> page = Fetcher.delete('https://example.com/resource/123', stealthy_headers=True, impersonate="chrome") ->>> page = Fetcher.delete('https://example.com/resource/123', proxy='http://username:password@localhost:8030') +from scrapling.fetchers import Fetcher +page = Fetcher.delete('https://example.com/resource/123') +page = Fetcher.delete('https://example.com/resource/123', stealthy_headers=True, impersonate="chrome") +page = Fetcher.delete('https://example.com/resource/123', proxy='http://username:password@localhost:8030') ``` And for asynchronous requests, it's a small adjustment ```python ->>> from scrapling.fetchers import AsyncFetcher ->>> page = await AsyncFetcher.delete('https://example.com/resource/123') ->>> page = await AsyncFetcher.delete('https://example.com/resource/123', stealthy_headers=True, impersonate="chrome") ->>> page = await AsyncFetcher.delete('https://example.com/resource/123', proxy='http://username:password@localhost:8030') +from scrapling.fetchers import AsyncFetcher +page = await AsyncFetcher.delete('https://example.com/resource/123') +page = await AsyncFetcher.delete('https://example.com/resource/123', stealthy_headers=True, impersonate="chrome") +page = await AsyncFetcher.delete('https://example.com/resource/123', proxy='http://username:password@localhost:8030') ``` ## Session Management diff --git a/agent-skill/Scrapling-Skill/references/fetching/stealthy.md b/agent-skill/Scrapling-Skill/references/fetching/stealthy.md index 8939c8d..ff4b2bf 100644 --- a/agent-skill/Scrapling-Skill/references/fetching/stealthy.md +++ b/agent-skill/Scrapling-Skill/references/fetching/stealthy.md @@ -6,7 +6,7 @@ You have one primary way to import this Fetcher, which is the same for all fetchers. ```python ->>> from scrapling.fetchers import StealthyFetcher +from scrapling.fetchers import StealthyFetcher ``` Check out how to configure the parsing options [here](choosing.md#parser-configuration-in-all-fetchers) diff --git a/agent-skill/Scrapling-Skill/references/parsing/adaptive.md b/agent-skill/Scrapling-Skill/references/parsing/adaptive.md index 61df382..25e2b79 100644 --- a/agent-skill/Scrapling-Skill/references/parsing/adaptive.md +++ b/agent-skill/Scrapling-Skill/references/parsing/adaptive.md @@ -68,22 +68,21 @@ To extract the Questions button from the old design, a selector like `#hmenus > Testing the same selector in both versions: ```python ->> from scrapling import Fetcher ->> selector = '#hmenus > div:nth-child(1) > ul > li:nth-child(1) > a' ->> old_url = "https://web.archive.org/web/20100102003420/http://stackoverflow.com/" ->> new_url = "https://stackoverflow.com/" ->> Fetcher.configure(adaptive = True, adaptive_domain='stackoverflow.com') ->> ->> page = Fetcher.get(old_url, timeout=30) ->> element1 = page.css(selector, auto_save=True)[0] ->> ->> # Same selector but used in the updated website ->> page = Fetcher.get(new_url) ->> element2 = page.css(selector, adaptive=True)[0] ->> ->> if element1.text == element2.text: +from scrapling import Fetcher +selector = '#hmenus > div:nth-child(1) > ul > li:nth-child(1) > a' +old_url = "https://web.archive.org/web/20100102003420/http://stackoverflow.com/" +new_url = "https://stackoverflow.com/" +Fetcher.configure(adaptive = True, adaptive_domain='stackoverflow.com') + +page = Fetcher.get(old_url, timeout=30) +element1 = page.css(selector, auto_save=True)[0] + +# Same selector but used in the updated website +page = Fetcher.get(new_url) +element2 = page.css(selector, adaptive=True)[0] + +if element1.text == element2.text: ... print('Scrapling found the same element in the old and new designs!') -'Scrapling found the same element in the old and new designs!' ``` The `adaptive_domain` argument is used here because Scrapling sees `archive.org` and `stackoverflow.com` as two different domains and would isolate their `adaptive` data. Passing `adaptive_domain` tells Scrapling to treat them as the same website for adaptive data storage. @@ -127,11 +126,11 @@ First, enable the `adaptive` feature by passing `adaptive=True` to the [Selector Examples: ```python ->>> from scrapling import Selector, Fetcher ->>> page = Selector(html_doc, adaptive=True) +from scrapling import Selector, Fetcher +page = Selector(html_doc, adaptive=True) # OR ->>> Fetcher.adaptive = True ->>> page = Fetcher.get('https://example.com') +Fetcher.adaptive = True +page = Fetcher.get('https://example.com') ``` When using the [Selector](main_classes.md#selector) class, pass the URL of the website with the `url` argument so Scrapling can separate the properties saved for each element by domain. @@ -159,11 +158,11 @@ Elements can be manually saved, retrieved, and relocated within the `adaptive` f Example of getting an element by text: ```python ->>> element = page.find_by_text('Tipping the Velvet', first_match=True) +element = page.find_by_text('Tipping the Velvet', first_match=True) ``` Save its unique properties using the `save` method. The identifier must be set manually (use a meaningful identifier): ```python ->>> page.save(element, 'my_special_element') +page.save(element, 'my_special_element') ``` Later, retrieve and relocate the element inside the page with `adaptive`: ```python diff --git a/agent-skill/Scrapling-Skill/references/parsing/main_classes.md b/agent-skill/Scrapling-Skill/references/parsing/main_classes.md index 3033e7c..0d54498 100644 --- a/agent-skill/Scrapling-Skill/references/parsing/main_classes.md +++ b/agent-skill/Scrapling-Skill/references/parsing/main_classes.md @@ -131,14 +131,14 @@ Getting the attributes of the element ``` Access a specific attribute with any of the following ```python ->>> article.attrib['class'] ->>> article.attrib.get('class') ->>> article['class'] # new in v0.3 +article.attrib['class'] +article.attrib.get('class') +article['class'] # new in v0.3 ``` Check if the attributes contain a specific attribute with any of the methods below ```python ->>> 'class' in article.attrib ->>> 'class' in article # new in v0.3 +'class' in article.attrib +'class' in article # new in v0.3 ``` Get the HTML content of the element ```python @@ -279,13 +279,13 @@ In the [Selector](#selector) class, all methods/properties that should return a Starting with v0.4, all selection methods consistently return [Selector](#selector)/[Selectors](#selectors) objects, even for text nodes and attribute values. Text nodes (selected via `::text`, `/text()`, `::attr()`, `/@attr`) are wrapped in [Selector](#selector) objects. These text node selectors have `tag` set to `"#text"`, and their `text` property returns the text value. You can still access the text value directly, and all other properties return empty/default values gracefully. ```python ->>> page.css('a::text') # -> Selectors (of text node Selectors) ->>> page.xpath('//a/text()') # -> Selectors ->>> page.css('a::text').get() # -> TextHandler (the first text value) ->>> page.css('a::text').getall() # -> TextHandlers (all text values) ->>> page.css('a::attr(href)') # -> Selectors ->>> page.xpath('//a/@href') # -> Selectors ->>> page.css('.price_color') # -> Selectors +page.css('a::text') # -> Selectors (of text node Selectors) +page.xpath('//a/text()') # -> Selectors +page.css('a::text').get() # -> TextHandler (the first text value) +page.css('a::text').getall() # -> TextHandlers (all text values) +page.css('a::attr(href)') # -> Selectors +page.xpath('//a/@href') # -> Selectors +page.css('.price_color') # -> Selectors ``` ### Data extraction methods diff --git a/agent-skill/Scrapling-Skill/references/parsing/selection.md b/agent-skill/Scrapling-Skill/references/parsing/selection.md index e74cbf1..d82a802 100644 --- a/agent-skill/Scrapling-Skill/references/parsing/selection.md +++ b/agent-skill/Scrapling-Skill/references/parsing/selection.md @@ -346,8 +346,8 @@ It filters all elements in the current page/element in the following order: ### Examples ```python ->>> from scrapling.fetchers import Fetcher ->>> page = Fetcher.get('https://quotes.toscrape.com/') +from scrapling.fetchers import Fetcher +page = Fetcher.get('https://quotes.toscrape.com/') ``` Find all elements with the tag name `div`. ```python diff --git a/docs/cli/interactive-shell.md b/docs/cli/interactive-shell.md index b897ce7..297aa35 100644 --- a/docs/cli/interactive-shell.md +++ b/docs/cli/interactive-shell.md @@ -43,25 +43,24 @@ Once launched, you'll see the Scrapling banner and can immediately start scrapin ```python # No imports needed - everything is ready! ->>> get('https://news.ycombinator.com') +get('https://news.ycombinator.com') ->>> # Explore the page structure ->>> page.css('a')[:5] # Look at first 5 links +# Explore the page structure +page.css('a')[:5] # Look at first 5 links ->>> # Refine your selectors ->>> stories = page.css('.titleline>a') ->>> len(stories) -30 +# Refine your selectors +stories = page.css('.titleline>a') +len(stories) # 30 ->>> # Extract specific data ->>> for story in stories[:3]: +# Extract specific data +for story in stories[:3]: ... title = story.text ... url = story['href'] ... print(f"{title}: {url}") ->>> # Try different approaches ->>> titles = page.css('.titleline>a::text') # Direct text extraction ->>> urls = page.css('.titleline>a::attr(href)') # Direct attribute extraction +# Try different approaches +titles = page.css('.titleline>a::text') # Direct text extraction +urls = page.css('.titleline>a::attr(href)') # Direct attribute extraction ``` ## Built-in Shortcuts @@ -86,12 +85,10 @@ The shell automatically tracks your requests and pages: The `page` and `response` commands are automatically updated with the last fetched page: ```python - >>> get('https://quotes.toscrape.com') - >>> # 'page' and 'response' both refer to the last fetched page - >>> page.url - 'https://quotes.toscrape.com' - >>> response.status # Same as page.status - 200 + get('https://quotes.toscrape.com') + # 'page' and 'response' both refer to the last fetched page + page.url # 'https://quotes.toscrape.com' + response.status # Prints 200; Same as page.status ``` - **Page History** @@ -99,20 +96,17 @@ The shell automatically tracks your requests and pages: The `pages` command keeps track of the last five pages (it's a `Selectors` object): ```python - >>> get('https://site1.com') - >>> get('https://site2.com') - >>> get('https://site3.com') + get('https://site1.com') + get('https://site2.com') + get('https://site3.com') - >>> # Access last 5 pages - >>> len(pages) # `Selectors` object with `page` history - 3 - >>> pages[0].url # First page in history - 'https://site1.com' - >>> pages[-1].url # Most recent page - 'https://site3.com' + # Access last 5 pages + len(pages) # `Selectors` object with `page` history -> 3 + pages[0].url # First page in history -> 'https://site1.com' + pages[-1].url # Most recent page -> 'https://site3.com' - >>> # Work with historical pages - >>> for i, old_page in enumerate(pages): + # Work with historical pages + for i, old_page in enumerate(pages): ... print(f"Page {i}: {old_page.url} - {old_page.status}") ``` @@ -123,8 +117,8 @@ The shell automatically tracks your requests and pages: View scraped pages in your browser: ```python ->>> get('https://quotes.toscrape.com') ->>> view(page) # Opens the page HTML in your default browser + get('https://quotes.toscrape.com') + view(page) # Opens the page HTML in your default browser ``` ### Curl Command Integration @@ -138,29 +132,24 @@ First, you need to copy a request as a curl command like the following: - **Convert Curl command to Request Object** ```python - >>> curl_cmd = '''curl 'https://scrapling.requestcatcher.com/post' \ + curl_cmd = '''curl 'https://scrapling.requestcatcher.com/post' \ ... -X POST \ ... -H 'Content-Type: application/json' \ ... -d '{"name": "test", "value": 123}' ''' - >>> request = uncurl(curl_cmd) - >>> request.method - 'post' - >>> request.url - 'https://scrapling.requestcatcher.com/post' - >>> request.headers - {'Content-Type': 'application/json'} + request = uncurl(curl_cmd) + request.method # -> 'post' + request.url # -> 'https://scrapling.requestcatcher.com/post' + request.headers # -> {'Content-Type': 'application/json'} ``` - **Execute Curl Command Directly** ```python - >>> # Convert and execute in one step - >>> curl2fetcher(curl_cmd) - >>> page.status - 200 - >>> page.json()['json'] - {'name': 'test', 'value': 123} + # Convert and execute in one step + curl2fetcher(curl_cmd) + page.status # -> 200 + page.json()['json'] # -> {'name': 'test', 'value': 123} ``` ### IPython Features @@ -168,17 +157,17 @@ First, you need to copy a request as a curl command like the following: The shell inherits all IPython capabilities: ```python ->>> # Magic commands ->>> %time page = get('https://example.com') # Time execution ->>> %history # Show command history ->>> %save filename.py 1-10 # Save commands 1-10 to file +# Magic commands +%time page = get('https://example.com') # Time execution +%history # Show command history +%save filename.py 1-10 # Save commands 1-10 to file ->>> # Tab completion works everywhere ->>> page.c # Shows: css, cookies, headers, etc. ->>> Fetcher. # Shows all Fetcher methods +# Tab completion works everywhere +page.c # Shows: css, cookies, headers, etc. +Fetcher. # Shows all Fetcher methods ->>> # Object inspection ->>> get? # Show get documentation +# Object inspection +get? # Show get documentation ``` ## Examples @@ -188,23 +177,23 @@ Here are a few examples generated via AI: #### E-commerce Data Collection ```python ->>> # Start with product listing page ->>> catalog = get('https://shop.example.com/products') +# Start with product listing page +catalog = get('https://shop.example.com/products') ->>> # Find product links ->>> product_links = catalog.css('.product-link::attr(href)') ->>> print(f"Found {len(product_links)} products") +# Find product links +product_links = catalog.css('.product-link::attr(href)') +print(f"Found {len(product_links)} products") ->>> # Sample a few products first ->>> for link in product_links[:3]: +# Sample a few products first +for link in product_links[:3]: ... product = get(f"https://shop.example.com{link}") ... name = product.css('.product-name::text').get('') ... price = product.css('.price::text').get('') ... print(f"{name}: {price}") ->>> # Scale up with sessions for efficiency ->>> from scrapling.fetchers import FetcherSession ->>> with FetcherSession() as session: +# Scale up with sessions for efficiency +from scrapling.fetchers import FetcherSession +with FetcherSession() as session: ... products = [] ... for link in product_links: ... product = session.get(f"https://shop.example.com{link}") diff --git a/docs/development/scrapling_custom_types.md b/docs/development/scrapling_custom_types.md index 2f638a9..ef67d27 100644 --- a/docs/development/scrapling_custom_types.md +++ b/docs/development/scrapling_custom_types.md @@ -4,13 +4,12 @@ ### All current types can be imported alone, like below ```python ->>> from scrapling.core.custom_types import TextHandler, AttributesHandler +from scrapling.core.custom_types import TextHandler, AttributesHandler ->>> somestring = TextHandler('{}') ->>> somestring.json() -'{}' ->>> somedict_1 = AttributesHandler({'a': 1}) ->>> somedict_2 = AttributesHandler(a=1) +somestring = TextHandler('{}') +somestring.json() # '{}' +somedict_1 = AttributesHandler({'a': 1}) +somedict_2 = AttributesHandler(a=1) ``` Note that `TextHandler` is a subclass of Python's `str`, so all standard operations/methods that work with Python strings will work. diff --git a/docs/fetching/choosing.md b/docs/fetching/choosing.md index b9f3ee4..b2b4bf2 100644 --- a/docs/fetching/choosing.md +++ b/docs/fetching/choosing.md @@ -31,23 +31,23 @@ In the following pages, we will talk about each one in detail. ## Parser configuration in all fetchers All fetchers share the same import method, as you will see in the upcoming pages ```python ->>> from scrapling.fetchers import Fetcher, AsyncFetcher, StealthyFetcher, DynamicFetcher +from scrapling.fetchers import Fetcher, AsyncFetcher, StealthyFetcher, DynamicFetcher ``` Then you use it right away without initializing like this, and it will use the default parser settings: ```python ->>> page = StealthyFetcher.fetch('https://example.com') +page = StealthyFetcher.fetch('https://example.com') ``` If you want to configure the parser ([Selector class](../parsing/main_classes.md#selector)) that will be used on the response before returning it for you, then do this first: ```python ->>> from scrapling.fetchers import Fetcher ->>> Fetcher.configure(adaptive=True, keep_comments=False, keep_cdata=False) # and the rest +from scrapling.fetchers import Fetcher +Fetcher.configure(adaptive=True, keep_comments=False, keep_cdata=False) # and the rest ``` or ```python ->>> from scrapling.fetchers import Fetcher ->>> Fetcher.adaptive=True ->>> Fetcher.keep_comments=False ->>> Fetcher.keep_cdata=False # and the rest +from scrapling.fetchers import Fetcher +Fetcher.adaptive=True +Fetcher.keep_comments=False +Fetcher.keep_cdata=False # and the rest ``` Then, continue your code as usual. @@ -65,19 +65,19 @@ If your use case requires a different configuration for each request/fetch, you ## Response Object The `Response` object is the same as the [Selector](../parsing/main_classes.md#selector) class, but it has additional details about the response, like response headers, status, cookies, etc., as shown below: ```python ->>> from scrapling.fetchers import Fetcher ->>> page = Fetcher.get('https://example.com') +from scrapling.fetchers import Fetcher +page = Fetcher.get('https://example.com') ->>> page.status # HTTP status code ->>> page.reason # Status message ->>> page.cookies # Response cookies as a dictionary ->>> page.headers # Response headers ->>> page.request_headers # Request headers ->>> page.history # Response history of redirections, if any ->>> page.body # Raw response body as bytes ->>> page.encoding # Response encoding ->>> page.meta # Response metadata dictionary (e.g., proxy used). Mainly helpful with the spiders system. ->>> page.captured_xhr # List of captured XHR/fetch responses (when capture_xhr is enabled on a browser session) +page.status # HTTP status code +page.reason # Status message +page.cookies # Response cookies as a dictionary +page.headers # Response headers +page.request_headers # Request headers +page.history # Response history of redirections, if any +page.body # Raw response body as bytes +page.encoding # Response encoding +page.meta # Response metadata dictionary (e.g., proxy used). Mainly helpful with the spiders system. +page.captured_xhr # List of captured XHR/fetch responses (when capture_xhr is enabled on a browser session) ``` All fetchers return the `Response` object. diff --git a/docs/fetching/dynamic.md b/docs/fetching/dynamic.md index 2e1f537..f987f4a 100644 --- a/docs/fetching/dynamic.md +++ b/docs/fetching/dynamic.md @@ -14,7 +14,7 @@ As we will explain later, to automate the page, you need some knowledge of [Play You have one primary way to import this Fetcher, which is the same for all fetchers. ```python ->>> from scrapling.fetchers import DynamicFetcher +from scrapling.fetchers import DynamicFetcher ``` Check out how to configure the parsing options [here](choosing.md#parser-configuration-in-all-fetchers) diff --git a/docs/fetching/static.md b/docs/fetching/static.md index 7d2a522..b6bcfef 100644 --- a/docs/fetching/static.md +++ b/docs/fetching/static.md @@ -12,7 +12,7 @@ The `Fetcher` class provides rapid and lightweight HTTP requests using the high- You have one primary way to import this Fetcher, which is the same for all fetchers. ```python ->>> from scrapling.fetchers import Fetcher +from scrapling.fetchers import Fetcher ``` Check out how to configure the parsing options [here](choosing.md#parser-configuration-in-all-fetchers) @@ -54,47 +54,47 @@ Examples are the best way to explain this: > Hence: `OPTIONS` and `HEAD` methods are not supported. #### GET ```python ->>> from scrapling.fetchers import Fetcher ->>> # Basic GET ->>> page = Fetcher.get('https://example.com') ->>> page = Fetcher.get('https://scrapling.requestcatcher.com/get', stealthy_headers=True) ->>> page = Fetcher.get('https://scrapling.requestcatcher.com/get', proxy='http://username:password@localhost:8030') ->>> # With parameters ->>> page = Fetcher.get('https://example.com/search', params={'q': 'query'}) ->>> ->>> # With headers ->>> page = Fetcher.get('https://example.com', headers={'User-Agent': 'Custom/1.0'}) ->>> # Basic HTTP authentication ->>> page = Fetcher.get("https://example.com", auth=("my_user", "password123")) ->>> # Browser impersonation ->>> page = Fetcher.get('https://example.com', impersonate='chrome') ->>> # HTTP/3 support ->>> page = Fetcher.get('https://example.com', http3=True) +from scrapling.fetchers import Fetcher +# Basic GET +page = Fetcher.get('https://example.com') +page = Fetcher.get('https://scrapling.requestcatcher.com/get', stealthy_headers=True) +page = Fetcher.get('https://scrapling.requestcatcher.com/get', proxy='http://username:password@localhost:8030') +# With parameters +page = Fetcher.get('https://example.com/search', params={'q': 'query'}) + +# With headers +page = Fetcher.get('https://example.com', headers={'User-Agent': 'Custom/1.0'}) +# Basic HTTP authentication +page = Fetcher.get("https://example.com", auth=("my_user", "password123")) +# Browser impersonation +page = Fetcher.get('https://example.com', impersonate='chrome') +# HTTP/3 support +page = Fetcher.get('https://example.com', http3=True) ``` And for asynchronous requests, it's a small adjustment ```python ->>> from scrapling.fetchers import AsyncFetcher ->>> # Basic GET ->>> page = await AsyncFetcher.get('https://example.com') ->>> page = await AsyncFetcher.get('https://scrapling.requestcatcher.com/get', stealthy_headers=True) ->>> page = await AsyncFetcher.get('https://scrapling.requestcatcher.com/get', proxy='http://username:password@localhost:8030') ->>> # With parameters ->>> page = await AsyncFetcher.get('https://example.com/search', params={'q': 'query'}) +from scrapling.fetchers import AsyncFetcher +# Basic GET +page = await AsyncFetcher.get('https://example.com') +page = await AsyncFetcher.get('https://scrapling.requestcatcher.com/get', stealthy_headers=True) +page = await AsyncFetcher.get('https://scrapling.requestcatcher.com/get', proxy='http://username:password@localhost:8030') +# With parameters + page = await AsyncFetcher.get('https://example.com/search', params={'q': 'query'}) >>> ->>> # With headers ->>> page = await AsyncFetcher.get('https://example.com', headers={'User-Agent': 'Custom/1.0'}) ->>> # Basic HTTP authentication ->>> page = await AsyncFetcher.get("https://example.com", auth=("my_user", "password123")) ->>> # Browser impersonation ->>> page = await AsyncFetcher.get('https://example.com', impersonate='chrome110') ->>> # HTTP/3 support ->>> page = await AsyncFetcher.get('https://example.com', http3=True) +# With headers +page = await AsyncFetcher.get('https://example.com', headers={'User-Agent': 'Custom/1.0'}) +# Basic HTTP authentication +page = await AsyncFetcher.get("https://example.com", auth=("my_user", "password123")) +# Browser impersonation +page = await AsyncFetcher.get('https://example.com', impersonate='chrome110') +# HTTP/3 support +page = await AsyncFetcher.get('https://example.com', http3=True) ``` Needless to say, the `page` object in all cases is [Response](choosing.md#response-object) object, which is a [Selector](../parsing/main_classes.md#selector) as we said, so you can use it directly ```python ->>> page.css('.something.something') +page.css('.something.something') ->>> page = Fetcher.get('https://api.github.com/events') +page = Fetcher.get('https://api.github.com/events') >>> page.json() [{'id': '', 'type': 'PushEvent', @@ -109,62 +109,62 @@ Needless to say, the `page` object in all cases is [Response](choosing.md#respon ``` #### POST ```python ->>> from scrapling.fetchers import Fetcher ->>> # Basic POST ->>> page = Fetcher.post('https://scrapling.requestcatcher.com/post', data={'key': 'value'}, params={'q': 'query'}) ->>> page = Fetcher.post('https://scrapling.requestcatcher.com/post', data={'key': 'value'}, stealthy_headers=True) ->>> page = Fetcher.post('https://scrapling.requestcatcher.com/post', data={'key': 'value'}, proxy='http://username:password@localhost:8030', impersonate="chrome") ->>> # Another example of form-encoded data ->>> page = Fetcher.post('https://example.com/submit', data={'username': 'user', 'password': 'pass'}, http3=True) ->>> # JSON data ->>> page = Fetcher.post('https://example.com/api', json={'key': 'value'}) +from scrapling.fetchers import Fetcher +# Basic POST +page = Fetcher.post('https://scrapling.requestcatcher.com/post', data={'key': 'value'}, params={'q': 'query'}) +page = Fetcher.post('https://scrapling.requestcatcher.com/post', data={'key': 'value'}, stealthy_headers=True) +page = Fetcher.post('https://scrapling.requestcatcher.com/post', data={'key': 'value'}, proxy='http://username:password@localhost:8030', impersonate="chrome") +# Another example of form-encoded data +page = Fetcher.post('https://example.com/submit', data={'username': 'user', 'password': 'pass'}, http3=True) +# JSON data +page = Fetcher.post('https://example.com/api', json={'key': 'value'}) ``` And for asynchronous requests, it's a small adjustment ```python ->>> from scrapling.fetchers import AsyncFetcher ->>> # Basic POST ->>> page = await AsyncFetcher.post('https://scrapling.requestcatcher.com/post', data={'key': 'value'}) ->>> page = await AsyncFetcher.post('https://scrapling.requestcatcher.com/post', data={'key': 'value'}, stealthy_headers=True) ->>> page = await AsyncFetcher.post('https://scrapling.requestcatcher.com/post', data={'key': 'value'}, proxy='http://username:password@localhost:8030', impersonate="chrome") ->>> # Another example of form-encoded data ->>> page = await AsyncFetcher.post('https://example.com/submit', data={'username': 'user', 'password': 'pass'}, http3=True) ->>> # JSON data ->>> page = await AsyncFetcher.post('https://example.com/api', json={'key': 'value'}) +from scrapling.fetchers import AsyncFetcher +# Basic POST +page = await AsyncFetcher.post('https://scrapling.requestcatcher.com/post', data={'key': 'value'}) +page = await AsyncFetcher.post('https://scrapling.requestcatcher.com/post', data={'key': 'value'}, stealthy_headers=True) +page = await AsyncFetcher.post('https://scrapling.requestcatcher.com/post', data={'key': 'value'}, proxy='http://username:password@localhost:8030', impersonate="chrome") +# Another example of form-encoded data +page = await AsyncFetcher.post('https://example.com/submit', data={'username': 'user', 'password': 'pass'}, http3=True) +# JSON data +page = await AsyncFetcher.post('https://example.com/api', json={'key': 'value'}) ``` #### PUT ```python ->>> from scrapling.fetchers import Fetcher ->>> # Basic PUT ->>> page = Fetcher.put('https://example.com/update', data={'status': 'updated'}) ->>> page = Fetcher.put('https://example.com/update', data={'status': 'updated'}, stealthy_headers=True, impersonate="chrome") ->>> page = Fetcher.put('https://example.com/update', data={'status': 'updated'}, proxy='http://username:password@localhost:8030') ->>> # Another example of form-encoded data ->>> page = Fetcher.put("https://scrapling.requestcatcher.com/put", data={'key': ['value1', 'value2']}) +from scrapling.fetchers import Fetcher +# Basic PUT +page = Fetcher.put('https://example.com/update', data={'status': 'updated'}) +page = Fetcher.put('https://example.com/update', data={'status': 'updated'}, stealthy_headers=True, impersonate="chrome") +page = Fetcher.put('https://example.com/update', data={'status': 'updated'}, proxy='http://username:password@localhost:8030') +# Another example of form-encoded data +page = Fetcher.put("https://scrapling.requestcatcher.com/put", data={'key': ['value1', 'value2']}) ``` And for asynchronous requests, it's a small adjustment ```python ->>> from scrapling.fetchers import AsyncFetcher ->>> # Basic PUT ->>> page = await AsyncFetcher.put('https://example.com/update', data={'status': 'updated'}) ->>> page = await AsyncFetcher.put('https://example.com/update', data={'status': 'updated'}, stealthy_headers=True, impersonate="chrome") ->>> page = await AsyncFetcher.put('https://example.com/update', data={'status': 'updated'}, proxy='http://username:password@localhost:8030') ->>> # Another example of form-encoded data ->>> page = await AsyncFetcher.put("https://scrapling.requestcatcher.com/put", data={'key': ['value1', 'value2']}) +from scrapling.fetchers import AsyncFetcher +# Basic PUT +page = await AsyncFetcher.put('https://example.com/update', data={'status': 'updated'}) +page = await AsyncFetcher.put('https://example.com/update', data={'status': 'updated'}, stealthy_headers=True, impersonate="chrome") +page = await AsyncFetcher.put('https://example.com/update', data={'status': 'updated'}, proxy='http://username:password@localhost:8030') +# Another example of form-encoded data +page = await AsyncFetcher.put("https://scrapling.requestcatcher.com/put", data={'key': ['value1', 'value2']}) ``` #### DELETE ```python ->>> from scrapling.fetchers import Fetcher ->>> page = Fetcher.delete('https://example.com/resource/123') ->>> page = Fetcher.delete('https://example.com/resource/123', stealthy_headers=True, impersonate="chrome") ->>> page = Fetcher.delete('https://example.com/resource/123', proxy='http://username:password@localhost:8030') +from scrapling.fetchers import Fetcher +page = Fetcher.delete('https://example.com/resource/123') +page = Fetcher.delete('https://example.com/resource/123', stealthy_headers=True, impersonate="chrome") +page = Fetcher.delete('https://example.com/resource/123', proxy='http://username:password@localhost:8030') ``` And for asynchronous requests, it's a small adjustment ```python ->>> from scrapling.fetchers import AsyncFetcher ->>> page = await AsyncFetcher.delete('https://example.com/resource/123') ->>> page = await AsyncFetcher.delete('https://example.com/resource/123', stealthy_headers=True, impersonate="chrome") ->>> page = await AsyncFetcher.delete('https://example.com/resource/123', proxy='http://username:password@localhost:8030') +from scrapling.fetchers import AsyncFetcher +page = await AsyncFetcher.delete('https://example.com/resource/123') +page = await AsyncFetcher.delete('https://example.com/resource/123', stealthy_headers=True, impersonate="chrome") +page = await AsyncFetcher.delete('https://example.com/resource/123', proxy='http://username:password@localhost:8030') ``` ## Session Management diff --git a/docs/fetching/stealthy.md b/docs/fetching/stealthy.md index 8daf0b0..2cff63d 100644 --- a/docs/fetching/stealthy.md +++ b/docs/fetching/stealthy.md @@ -15,7 +15,7 @@ As with [DynamicFetcher](dynamic.md#introduction), you will need some knowledge You have one primary way to import this Fetcher, which is the same for all fetchers. ```python ->>> from scrapling.fetchers import StealthyFetcher + from scrapling.fetchers import StealthyFetcher ``` Check out how to configure the parsing options [here](choosing.md#parser-configuration-in-all-fetchers) diff --git a/docs/overview.md b/docs/overview.md index c62c7c3..5a72b95 100644 --- a/docs/overview.md +++ b/docs/overview.md @@ -263,19 +263,19 @@ page = Fetcher.get('https://scrapling.requestcatcher.com/get', impersonate="chro ``` With that out of the way, here's how to do all HTTP methods: ```python ->>> from scrapling.fetchers import Fetcher ->>> page = Fetcher.get('https://scrapling.requestcatcher.com/get', stealthy_headers=True) ->>> page = Fetcher.post('https://scrapling.requestcatcher.com/post', data={'key': 'value'}, proxy='http://username:password@localhost:8030') ->>> page = Fetcher.put('https://scrapling.requestcatcher.com/put', data={'key': 'value'}) ->>> page = Fetcher.delete('https://scrapling.requestcatcher.com/delete') +from scrapling.fetchers import Fetcher +page = Fetcher.get('https://scrapling.requestcatcher.com/get', stealthy_headers=True) +page = Fetcher.post('https://scrapling.requestcatcher.com/post', data={'key': 'value'}, proxy='http://username:password@localhost:8030') +page = Fetcher.put('https://scrapling.requestcatcher.com/put', data={'key': 'value'}) +page = Fetcher.delete('https://scrapling.requestcatcher.com/delete') ``` For Async requests, you will replace the import like below: ```python ->>> from scrapling.fetchers import AsyncFetcher ->>> page = await AsyncFetcher.get('https://scrapling.requestcatcher.com/get', stealthy_headers=True) ->>> page = await AsyncFetcher.post('https://scrapling.requestcatcher.com/post', data={'key': 'value'}, proxy='http://username:password@localhost:8030') ->>> page = await AsyncFetcher.put('https://scrapling.requestcatcher.com/put', data={'key': 'value'}) ->>> page = await AsyncFetcher.delete('https://scrapling.requestcatcher.com/delete') +from scrapling.fetchers import AsyncFetcher +page = await AsyncFetcher.get('https://scrapling.requestcatcher.com/get', stealthy_headers=True) +page = await AsyncFetcher.post('https://scrapling.requestcatcher.com/post', data={'key': 'value'}, proxy='http://username:password@localhost:8030') +page = await AsyncFetcher.put('https://scrapling.requestcatcher.com/put', data={'key': 'value'}) +page = await AsyncFetcher.delete('https://scrapling.requestcatcher.com/delete') ``` !!! note "Notes:" @@ -291,14 +291,13 @@ We have you covered if you deal with dynamic websites like most today! The `DynamicFetcher` class (formerly `PlayWrightFetcher`) offers many options for fetching and loading web pages using Chromium-based browsers. ```python ->>> from scrapling.fetchers import DynamicFetcher ->>> page = DynamicFetcher.fetch('https://www.google.com/search?q=%22Scrapling%22', disable_resources=True) # Vanilla Playwright option ->>> page.css("#search a::attr(href)").get() -'https://github.com/D4Vinci/Scrapling' ->>> # The async version of fetch ->>> page = await DynamicFetcher.async_fetch('https://www.google.com/search?q=%22Scrapling%22', disable_resources=True) ->>> page.css("#search a::attr(href)").get() -'https://github.com/D4Vinci/Scrapling' +from scrapling.fetchers import DynamicFetcher +page = DynamicFetcher.fetch('https://www.google.com/search?q=%22Scrapling%22', disable_resources=True) # Vanilla Playwright option +page.css("#search a::attr(href)").get() # -> 'https://github.com/D4Vinci/Scrapling' + +# The async version of fetch +page = await DynamicFetcher.async_fetch('https://www.google.com/search?q=%22Scrapling%22', disable_resources=True) +page.css("#search a::attr(href)").get() # -> 'https://github.com/D4Vinci/Scrapling' ``` It's built on top of [Playwright](https://playwright.dev/python/), and it's currently providing two main run options that can be mixed as you want: @@ -323,18 +322,17 @@ Some of the things it does: 6. and other anti-protection options... ```python ->>> from scrapling.fetchers import StealthyFetcher ->>> page = StealthyFetcher.fetch('https://www.browserscan.net/bot-detection') # Running headless by default ->>> page.status == 200 -True ->>> page = StealthyFetcher.fetch('https://nopecha.com/demo/cloudflare', solve_cloudflare=True) # Solve Cloudflare captcha automatically if presented ->>> page.status == 200 -True ->>> page = StealthyFetcher.fetch('https://www.browserscan.net/bot-detection', humanize=True, os_randomize=True) # and the rest of arguments... ->>> # The async version of fetch ->>> page = await StealthyFetcher.async_fetch('https://www.browserscan.net/bot-detection') ->>> page.status == 200 -True +from scrapling.fetchers import StealthyFetcher +page = StealthyFetcher.fetch('https://www.browserscan.net/bot-detection') # Running headless by default +page.status == 200 # -> True + +page = StealthyFetcher.fetch('https://nopecha.com/demo/cloudflare', solve_cloudflare=True) # Solve Cloudflare captcha automatically if presented +page.status == 200 # -> True + +page = StealthyFetcher.fetch('https://www.browserscan.net/bot-detection', humanize=True, os_randomize=True) # and the rest of arguments... +# The async version of fetch +page = await StealthyFetcher.async_fetch('https://www.browserscan.net/bot-detection') +page.status == 200 # -> True ``` Again, this is just the tip of the iceberg with this fetcher. Check out the rest from [here](fetching/stealthy.md) for all details and the complete list of arguments. diff --git a/docs/parsing/adaptive.md b/docs/parsing/adaptive.md index 23dcaf3..b2ba6cf 100644 --- a/docs/parsing/adaptive.md +++ b/docs/parsing/adaptive.md @@ -76,22 +76,19 @@ If I want to extract the Questions button from the old design, I can use a selec Now, let's test the same selector in both versions ```python ->> from scrapling import Fetcher ->> selector = '#hmenus > div:nth-child(1) > ul > li:nth-child(1) > a' ->> old_url = "https://web.archive.org/web/20100102003420/http://stackoverflow.com/" ->> new_url = "https://stackoverflow.com/" ->> Fetcher.configure(adaptive = True, adaptive_domain='stackoverflow.com') ->> ->> page = Fetcher.get(old_url, timeout=30) ->> element1 = page.css(selector, auto_save=True)[0] ->> ->> # Same selector but used in the updated website ->> page = Fetcher.get(new_url) ->> element2 = page.css(selector, adaptive=True)[0] ->> ->> if element1.text == element2.text: -... print('Scrapling found the same element in the old and new designs!') -'Scrapling found the same element in the old and new designs!' +from scrapling import Fetcher +selector = '#hmenus > div:nth-child(1) > ul > li:nth-child(1) > a' +old_url = "https://web.archive.org/web/20100102003420/http://stackoverflow.com/" +new_url = "https://stackoverflow.com/" +Fetcher.configure(adaptive = True, adaptive_domain='stackoverflow.com') +page = Fetcher.get(old_url, timeout=30) +element1 = page.css(selector, auto_save=True)[0] +# Same selector but used in the updated website +page = Fetcher.get(new_url) +element2 = page.css(selector, adaptive=True)[0] + +if element1.text == element2.text: + print('Scrapling found the same element in the old and new designs!') # Spoiler alert: it does! ``` Note that I introduced a new argument called `adaptive_domain`. This is because, for Scrapling, these are two different domains (`archive.org` and `stackoverflow.com`), so Scrapling will isolate their `adaptive` data. To inform Scrapling that they are the same website, we must pass the custom domain we wish to use while saving `adaptive` data for both, ensuring Scrapling doesn't isolate them. @@ -141,11 +138,11 @@ First, you must enable the `adaptive` feature by passing `adaptive=True` to the Examples: ```python ->>> from scrapling import Selector, Fetcher ->>> page = Selector(html_doc, adaptive=True) +from scrapling import Selector, Fetcher +page = Selector(html_doc, adaptive=True) # OR ->>> Fetcher.adaptive = True ->>> page = Fetcher.get('https://example.com') +Fetcher.adaptive = True +page = Fetcher.get('https://example.com') ``` If you are using the [Selector](main_classes.md#selector) class, you need to pass the url of the website you are using with the argument `url` so Scrapling can separate the properties saved for each element by domain. @@ -175,11 +172,11 @@ You manually save and retrieve an element, then relocate it, which all happens w First, let's say you got an element like this by text: ```python ->>> element = page.find_by_text('Tipping the Velvet', first_match=True) +element = page.find_by_text('Tipping the Velvet', first_match=True) ``` You can save its unique properties using the `save` method, as shown below, but you must set the identifier yourself. For this example, I chose `my_special_element` as an identifier, but it's best to use a meaningful identifier in your code for the same reason you use meaningful variable names :) ```python ->>> page.save(element, 'my_special_element') +page.save(element, 'my_special_element') ``` Now, later, when you want to retrieve it and relocate it inside the page with `adaptive`, it would be like this ```python diff --git a/docs/parsing/main_classes.md b/docs/parsing/main_classes.md index 6d2cac0..d41ff40 100644 --- a/docs/parsing/main_classes.md +++ b/docs/parsing/main_classes.md @@ -140,14 +140,14 @@ Getting the attributes of the element ``` Access a specific attribute with any of the following ```python ->>> article.attrib['class'] ->>> article.attrib.get('class') ->>> article['class'] # new in v0.3 +article.attrib['class'] +article.attrib.get('class') +article['class'] # new in v0.3 ``` Check if the attributes contain a specific attribute with any of the methods below ```python ->>> 'class' in article.attrib ->>> 'class' in article # new in v0.3 +'class' in article.attrib +'class' in article # new in v0.3 ``` Get the HTML content of the element ```python @@ -292,13 +292,13 @@ In the [Selector](#selector) class, all methods/properties that should return a Starting with v0.4, all selection methods consistently return [Selector](#selector)/[Selectors](#selectors) objects, even for text nodes and attribute values. Text nodes (selected via `::text`, `/text()`, `::attr()`, `/@attr`) are wrapped in [Selector](#selector) objects. These text node selectors have `tag` set to `"#text"`, and their `text` property returns the text value. You can still access the text value directly, and all other properties return empty/default values gracefully. ```python ->>> page.css('a::text') # -> Selectors (of text node Selectors) ->>> page.xpath('//a/text()') # -> Selectors ->>> page.css('a::text').get() # -> TextHandler (the first text value) ->>> page.css('a::text').getall() # -> TextHandlers (all text values) ->>> page.css('a::attr(href)') # -> Selectors ->>> page.xpath('//a/@href') # -> Selectors ->>> page.css('.price_color') # -> Selectors +page.css('a::text') # -> Selectors (of text node Selectors) +page.xpath('//a/text()') # -> Selectors +page.css('a::text').get() # -> TextHandler (the first text value) +page.css('a::text').getall() # -> TextHandlers (all text values) +page.css('a::attr(href)') # -> Selectors +page.xpath('//a/@href') # -> Selectors +page.css('.price_color') # -> Selectors ``` ### Data extraction methods diff --git a/docs/parsing/selection.md b/docs/parsing/selection.md index 5311391..aef8f3b 100644 --- a/docs/parsing/selection.md +++ b/docs/parsing/selection.md @@ -362,8 +362,8 @@ Check examples to clear any confusion :) ### Examples ```python ->>> from scrapling.fetchers import Fetcher ->>> page = Fetcher.get('https://quotes.toscrape.com/') +from scrapling.fetchers import Fetcher +page = Fetcher.get('https://quotes.toscrape.com/') ``` Find all elements with the tag name `div`. ```python