docs: adding separate doc for external browser
This commit is contained in:
@@ -302,39 +302,3 @@ Use DynamicFetcher when:
|
|||||||
- Want flexible stealth options
|
- Want flexible stealth options
|
||||||
|
|
||||||
If you want more stealth and control without much config, check out the [StealthyFetcher](stealthy.md).
|
If you want more stealth and control without much config, check out the [StealthyFetcher](stealthy.md).
|
||||||
|
|
||||||
## External Cloud Browser Version
|
|
||||||
|
|
||||||
If you have issues with the browser installation, such as resource management, we recommend you try the Cloud Browser from [Scrapeless](https://www.scrapeless.com/en/product/scraping-browser) for free!
|
|
||||||
|
|
||||||
The usage is straightforward: create an account and [get your API key](https://docs.scrapeless.com/en/scraping-browser/quickstart/getting-started/), then pass it to the `DynamicSession` like this:
|
|
||||||
|
|
||||||
```python
|
|
||||||
from urllib.parse import urlencode
|
|
||||||
|
|
||||||
from scrapling.fetchers import DynamicSession
|
|
||||||
|
|
||||||
# Configure your browser session
|
|
||||||
config = {
|
|
||||||
"token": "YOUR_API_KEY",
|
|
||||||
"sessionName": "scrapling-session",
|
|
||||||
"sessionTTL": "300", # 5 minutes
|
|
||||||
"proxyCountry": "ANY",
|
|
||||||
"sessionRecording": "false",
|
|
||||||
}
|
|
||||||
|
|
||||||
# Build WebSocket URL
|
|
||||||
ws_endpoint = f"wss://browser.scrapeless.com/api/v2/browser?{urlencode(config)}"
|
|
||||||
print('Connecting to Scrapeless...')
|
|
||||||
|
|
||||||
with DynamicSession(cdp_url=ws_endpoint, disable_resources=True) as s:
|
|
||||||
print("Connected!")
|
|
||||||
page = s.fetch("https://httpbin.org/headers", network_idle=True)
|
|
||||||
print(f"Page loaded, content length: {len(page.body)}")
|
|
||||||
print(page.json())
|
|
||||||
```
|
|
||||||
The `DynamicSession` class instance will work as usual, so no further explanation is needed.
|
|
||||||
|
|
||||||
However, the Scrapeless Cloud Browser can be configured with proxy options, like the proxy country in the config above, [custom fingerprint](https://docs.scrapeless.com/en/scraping-browser/features/advanced-privacy-anti-detection/custom-fingerprint/) configuration, [captcha solving](https://docs.scrapeless.com/en/scraping-browser/features/advanced-privacy-anti-detection/supported-captchas/), and more.
|
|
||||||
|
|
||||||
Check out the [Scrapeless's browser documentation](https://docs.scrapeless.com/en/scraping-browser/quickstart/introduction/) for more details.
|
|
||||||
@@ -0,0 +1,35 @@
|
|||||||
|
## External Cloud Browser Version
|
||||||
|
|
||||||
|
If you have issues with the browser installation, such as resource management, we recommend you try the Cloud Browser from [Scrapeless](https://www.scrapeless.com/en/product/scraping-browser) for free!
|
||||||
|
|
||||||
|
The usage is straightforward: create an account and [get your API key](https://docs.scrapeless.com/en/scraping-browser/quickstart/getting-started/), then pass it to the `DynamicSession` like this:
|
||||||
|
|
||||||
|
```python
|
||||||
|
from urllib.parse import urlencode
|
||||||
|
|
||||||
|
from scrapling.fetchers import DynamicSession
|
||||||
|
|
||||||
|
# Configure your browser session
|
||||||
|
config = {
|
||||||
|
"token": "YOUR_API_KEY",
|
||||||
|
"sessionName": "scrapling-session",
|
||||||
|
"sessionTTL": "300", # 5 minutes
|
||||||
|
"proxyCountry": "ANY",
|
||||||
|
"sessionRecording": "false",
|
||||||
|
}
|
||||||
|
|
||||||
|
# Build WebSocket URL
|
||||||
|
ws_endpoint = f"wss://browser.scrapeless.com/api/v2/browser?{urlencode(config)}"
|
||||||
|
print('Connecting to Scrapeless...')
|
||||||
|
|
||||||
|
with DynamicSession(cdp_url=ws_endpoint, disable_resources=True) as s:
|
||||||
|
print("Connected!")
|
||||||
|
page = s.fetch("https://httpbin.org/headers", network_idle=True)
|
||||||
|
print(f"Page loaded, content length: {len(page.body)}")
|
||||||
|
print(page.json())
|
||||||
|
```
|
||||||
|
The `DynamicSession` class instance will work as usual, so no further explanation is needed.
|
||||||
|
|
||||||
|
However, the Scrapeless Cloud Browser can be configured with proxy options, like the proxy country in the config above, [custom fingerprint](https://docs.scrapeless.com/en/scraping-browser/features/advanced-privacy-anti-detection/custom-fingerprint/) configuration, [captcha solving](https://docs.scrapeless.com/en/scraping-browser/features/advanced-privacy-anti-detection/supported-captchas/), and more.
|
||||||
|
|
||||||
|
Check out the [Scrapeless's browser documentation](https://docs.scrapeless.com/en/scraping-browser/quickstart/introduction/) for more details.
|
||||||
@@ -83,6 +83,7 @@ nav:
|
|||||||
- Tutorials:
|
- Tutorials:
|
||||||
- A Free Alternative to AI for Robust Web Scraping: tutorials/replacing_ai.md
|
- A Free Alternative to AI for Robust Web Scraping: tutorials/replacing_ai.md
|
||||||
- Migrating from BeautifulSoup: tutorials/migrating_from_beautifulsoup.md
|
- Migrating from BeautifulSoup: tutorials/migrating_from_beautifulsoup.md
|
||||||
|
- Using Scrapeless browser: tutorials/external.md
|
||||||
# - Migrating from AutoScraper: tutorials/migrating_from_autoscraper.md
|
# - Migrating from AutoScraper: tutorials/migrating_from_autoscraper.md
|
||||||
- Development:
|
- Development:
|
||||||
- API Reference:
|
- API Reference:
|
||||||
|
|||||||
Reference in New Issue
Block a user