Merge branch 'dev' into test/edge-cases-filter-ancestors-find-similar
This commit is contained in:
@@ -38,17 +38,17 @@
|
||||
</p>
|
||||
|
||||
<p align="center">
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/parsing/selection/"><strong>Selection methods</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/parsing/selection.html"><strong>Selection methods</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/fetching/choosing/"><strong>Fetchers</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/fetching/choosing.html"><strong>Fetchers</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/spiders/architecture.html"><strong>Spiders</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/spiders/proxy-blocking.html"><strong>Proxy Rotation</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/cli/overview/"><strong>CLI</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/cli/overview.html"><strong>CLI</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/ai/mcp-server/"><strong>MCP</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/ai/mcp-server.html"><strong>MCP</strong></a>
|
||||
</p>
|
||||
|
||||
Scrapling is an adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl.
|
||||
@@ -163,6 +163,16 @@ MySpider().start()
|
||||
Read a full review of <a href="https://substack.thewebscraping.club/p/scrapling-hands-on-guide?utm_source=github&utm_medium=repo&utm_campaign=scrapling" target="_blank">Scrapling on The Web Scraping Club</a> (Nov 2025), the #1 newsletter dedicated to Web Scraping.
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td width="200">
|
||||
<a href="https://proxy-seller.com/?partner=CU9CAA5TBYFFT2" target="_blank" title="Proxy-Seller provides reliable proxy infrastructure for Web Scraping">
|
||||
<img src="https://raw.githubusercontent.com/D4Vinci/Scrapling/main/images/ProxySeller.png">
|
||||
</a>
|
||||
</td>
|
||||
<td>
|
||||
<a href="https://proxy-seller.com/?partner=CU9CAA5TBYFFT2" target="_blank">Proxy-Seller</a> provides reliable proxy infrastructure for web scraping, offering IPv4, IPv6, ISP, Residential, and Mobile proxies with stable performance, broad geo coverage, and flexible plans for business-scale data collection.
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
<i><sub>Do you want to show your ad here? Click [here](https://github.com/sponsors/D4Vinci/sponsorships?tier_id=586646)</sub></i>
|
||||
|
||||
+14
-4
@@ -34,17 +34,17 @@
|
||||
</p>
|
||||
|
||||
<p align="center">
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/parsing/selection/"><strong>طرق الاختيار</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/parsing/selection.html"><strong>طرق الاختيار</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/fetching/choosing/"><strong>اختيار Fetcher</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/fetching/choosing.html"><strong>اختيار Fetcher</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/spiders/architecture.html"><strong>العناكب</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/spiders/proxy-blocking.html"><strong>تدوير البروكسي</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/cli/overview/"><strong>واجهة سطر الأوامر</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/cli/overview.html"><strong>واجهة سطر الأوامر</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/ai/mcp-server/"><strong>وضع MCP</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/ai/mcp-server.html"><strong>وضع MCP</strong></a>
|
||||
</p>
|
||||
|
||||
Scrapling هو إطار عمل تكيفي لـ Web Scraping يتعامل مع كل شيء من طلب واحد إلى زحف كامل النطاق.
|
||||
@@ -159,6 +159,16 @@ MySpider().start()
|
||||
اقرأ مراجعة كاملة عن <a href="https://substack.thewebscraping.club/p/scrapling-hands-on-guide?utm_source=github&utm_medium=repo&utm_campaign=scrapling" target="_blank">Scrapling على The Web Scraping Club</a> (نوفمبر 2025)، النشرة الإخبارية الأولى المخصصة لكشط الويب.
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td width="200">
|
||||
<a href="https://proxy-seller.com/?partner=CU9CAA5TBYFFT2" target="_blank" title="Proxy-Seller provides reliable proxy infrastructure for Web Scraping">
|
||||
<img src="https://raw.githubusercontent.com/D4Vinci/Scrapling/main/images/ProxySeller.png">
|
||||
</a>
|
||||
</td>
|
||||
<td>
|
||||
<a href="https://proxy-seller.com/?partner=CU9CAA5TBYFFT2" target="_blank">Proxy-Seller</a> يوفر بنية تحتية موثوقة للبروكسي لكشط الويب، بما في ذلك بروكسيات IPv4 وIPv6 وISP والسكنية والمحمولة مع أداء مستقر وتغطية جغرافية واسعة وخطط مرنة لجمع البيانات على نطاق الأعمال.
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
<i><sub>هل تريد عرض إعلانك هنا؟ انقر [هنا](https://github.com/sponsors/D4Vinci/sponsorships?tier_id=586646)</sub></i>
|
||||
|
||||
+14
-4
@@ -34,17 +34,17 @@
|
||||
</p>
|
||||
|
||||
<p align="center">
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/parsing/selection/"><strong>选择方法</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/parsing/selection.html"><strong>选择方法</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/fetching/choosing/"><strong>选择Fetcher</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/fetching/choosing.html"><strong>选择 Fetcher</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/spiders/architecture.html"><strong>爬虫</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/spiders/proxy-blocking.html"><strong>代理轮换</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/cli/overview/"><strong>CLI</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/cli/overview.html"><strong>CLI</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/ai/mcp-server/"><strong>MCP模式</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/ai/mcp-server.html"><strong>MCP 模式</strong></a>
|
||||
</p>
|
||||
|
||||
Scrapling 是一个自适应 Web Scraping 框架,能处理从单个请求到大规模爬取的一切需求。
|
||||
@@ -159,6 +159,16 @@ MySpider().start()
|
||||
阅读 <a href="https://substack.thewebscraping.club/p/scrapling-hands-on-guide?utm_source=github&utm_medium=repo&utm_campaign=scrapling" target="_blank">The Web Scraping Club 上关于 Scrapling 的完整评测</a>(2025 年 11 月),这是排名第一的网页抓取专业通讯。
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td width="200">
|
||||
<a href="https://proxy-seller.com/?partner=CU9CAA5TBYFFT2" target="_blank" title="Proxy-Seller provides reliable proxy infrastructure for Web Scraping">
|
||||
<img src="https://raw.githubusercontent.com/D4Vinci/Scrapling/main/images/ProxySeller.png">
|
||||
</a>
|
||||
</td>
|
||||
<td>
|
||||
<a href="https://proxy-seller.com/?partner=CU9CAA5TBYFFT2" target="_blank">Proxy-Seller</a> 提供可靠的网页抓取代理基础设施,包括 IPv4、IPv6、ISP、住宅和移动代理,具备稳定性能、广泛的地理覆盖和灵活的企业级数据采集方案。
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
<i><sub>想在这里展示您的广告吗?点击 [这里](https://github.com/sponsors/D4Vinci/sponsorships?tier_id=586646)</sub></i>
|
||||
|
||||
+14
-4
@@ -34,17 +34,17 @@
|
||||
</p>
|
||||
|
||||
<p align="center">
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/parsing/selection/"><strong>Auswahlmethoden</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/parsing/selection.html"><strong>Auswahlmethoden</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/fetching/choosing/"><strong>Einen Fetcher wählen</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/fetching/choosing.html"><strong>Einen Fetcher wählen</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/spiders/architecture.html"><strong>Spiders</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/spiders/proxy-blocking.html"><strong>Proxy-Rotation</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/cli/overview/"><strong>CLI</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/cli/overview.html"><strong>CLI</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/ai/mcp-server/"><strong>MCP-Modus</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/ai/mcp-server.html"><strong>MCP-Modus</strong></a>
|
||||
</p>
|
||||
|
||||
Scrapling ist ein adaptives Web-Scraping-Framework, das alles abdeckt -- von einer einzelnen Anfrage bis hin zu einem umfassenden Crawl.
|
||||
@@ -159,6 +159,16 @@ MySpider().start()
|
||||
Lesen Sie eine vollständige Rezension von <a href="https://substack.thewebscraping.club/p/scrapling-hands-on-guide?utm_source=github&utm_medium=repo&utm_campaign=scrapling" target="_blank">Scrapling auf The Web Scraping Club</a> (Nov. 2025), dem führenden Newsletter für Web Scraping.
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td width="200">
|
||||
<a href="https://proxy-seller.com/?partner=CU9CAA5TBYFFT2" target="_blank" title="Proxy-Seller provides reliable proxy infrastructure for Web Scraping">
|
||||
<img src="https://raw.githubusercontent.com/D4Vinci/Scrapling/main/images/ProxySeller.png">
|
||||
</a>
|
||||
</td>
|
||||
<td>
|
||||
<a href="https://proxy-seller.com/?partner=CU9CAA5TBYFFT2" target="_blank">Proxy-Seller</a> bietet zuverlässige Proxy-Infrastruktur für Web Scraping mit IPv4-, IPv6-, ISP-, Residential- und Mobile-Proxys – stabile Leistung, breite geografische Abdeckung und flexible Tarife für die Datenerfassung im Unternehmensmaßstab.
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
<i><sub>Möchten Sie Ihre Anzeige hier zeigen? Klicken Sie [hier](https://github.com/sponsors/D4Vinci/sponsorships?tier_id=586646)</sub></i>
|
||||
|
||||
+14
-4
@@ -34,17 +34,17 @@
|
||||
</p>
|
||||
|
||||
<p align="center">
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/parsing/selection/"><strong>Métodos de selección</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/parsing/selection.html"><strong>Métodos de selección</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/fetching/choosing/"><strong>Elegir un fetcher</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/fetching/choosing.html"><strong>Elegir un fetcher</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/spiders/architecture.html"><strong>Spiders</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/spiders/proxy-blocking.html"><strong>Rotación de proxy</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/cli/overview/"><strong>CLI</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/cli/overview.html"><strong>CLI</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/ai/mcp-server/"><strong>Modo MCP</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/ai/mcp-server.html"><strong>Modo MCP</strong></a>
|
||||
</p>
|
||||
|
||||
Scrapling es un framework de Web Scraping adaptativo que se encarga de todo, desde una sola solicitud hasta un rastreo a gran escala.
|
||||
@@ -159,6 +159,16 @@ MySpider().start()
|
||||
Lee una reseña completa de <a href="https://substack.thewebscraping.club/p/scrapling-hands-on-guide?utm_source=github&utm_medium=repo&utm_campaign=scrapling" target="_blank">Scrapling en The Web Scraping Club</a> (nov. 2025), el boletín número uno dedicado al Web Scraping.
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td width="200">
|
||||
<a href="https://proxy-seller.com/?partner=CU9CAA5TBYFFT2" target="_blank" title="Proxy-Seller provides reliable proxy infrastructure for Web Scraping">
|
||||
<img src="https://raw.githubusercontent.com/D4Vinci/Scrapling/main/images/ProxySeller.png">
|
||||
</a>
|
||||
</td>
|
||||
<td>
|
||||
<a href="https://proxy-seller.com/?partner=CU9CAA5TBYFFT2" target="_blank">Proxy-Seller</a> ofrece una infraestructura de proxy fiable para web scraping, con proxies IPv4, IPv6, ISP, residenciales y móviles con rendimiento estable, amplia cobertura geográfica y planes flexibles para la recopilación de datos a escala empresarial.
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
<i><sub>¿Quieres mostrar tu anuncio aquí? Haz clic [aquí](https://github.com/sponsors/D4Vinci/sponsorships?tier_id=586646)</sub></i>
|
||||
|
||||
+14
-4
@@ -34,17 +34,17 @@
|
||||
</p>
|
||||
|
||||
<p align="center">
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/parsing/selection/"><strong>Méthodes de sélection</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/parsing/selection.html"><strong>Méthodes de sélection</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/fetching/choosing/"><strong>Fetchers</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/fetching/choosing.html"><strong>Fetchers</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/spiders/architecture.html"><strong>Spiders</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/spiders/proxy-blocking.html"><strong>Rotation de proxy</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/cli/overview/"><strong>CLI</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/cli/overview.html"><strong>CLI</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/ai/mcp-server/"><strong>MCP</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/ai/mcp-server.html"><strong>MCP</strong></a>
|
||||
</p>
|
||||
|
||||
Scrapling est un framework de Web Scraping adaptatif qui gère tout, d'une simple requête à un crawl à grande échelle.
|
||||
@@ -159,6 +159,16 @@ MySpider().start()
|
||||
Lisez une critique complète de <a href="https://substack.thewebscraping.club/p/scrapling-hands-on-guide?utm_source=github&utm_medium=repo&utm_campaign=scrapling" target="_blank">Scrapling sur The Web Scraping Club</a> (nov. 2025), la newsletter n°1 dédiée au Web Scraping.
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td width="200">
|
||||
<a href="https://proxy-seller.com/?partner=CU9CAA5TBYFFT2" target="_blank" title="Proxy-Seller provides reliable proxy infrastructure for Web Scraping">
|
||||
<img src="https://raw.githubusercontent.com/D4Vinci/Scrapling/main/images/ProxySeller.png">
|
||||
</a>
|
||||
</td>
|
||||
<td>
|
||||
<a href="https://proxy-seller.com/?partner=CU9CAA5TBYFFT2" target="_blank">Proxy-Seller</a> fournit une infrastructure proxy fiable pour le web scraping, avec des proxys IPv4, IPv6, ISP, résidentiels et mobiles offrant des performances stables, une large couverture géographique et des plans flexibles pour la collecte de données à l'échelle entreprise.
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
<i><sub>Vous souhaitez afficher votre publicité ici ? Cliquez [ici](https://github.com/sponsors/D4Vinci/sponsorships?tier_id=586646)</sub></i>
|
||||
|
||||
+14
-4
@@ -34,17 +34,17 @@
|
||||
</p>
|
||||
|
||||
<p align="center">
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/parsing/selection/"><strong>選択メソッド</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/parsing/selection.html"><strong>選択メソッド</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/fetching/choosing/"><strong>Fetcherの選び方</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/fetching/choosing.html"><strong>Fetcher の選び方</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/spiders/architecture.html"><strong>スパイダー</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/spiders/proxy-blocking.html"><strong>プロキシローテーション</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/cli/overview/"><strong>CLI</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/cli/overview.html"><strong>CLI</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/ai/mcp-server/"><strong>MCPモード</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/ai/mcp-server.html"><strong>MCP モード</strong></a>
|
||||
</p>
|
||||
|
||||
Scrapling は、単一のリクエストから本格的なクロールまですべてを処理する適応型 Web Scraping フレームワークです。
|
||||
@@ -159,6 +159,16 @@ MySpider().start()
|
||||
<a href="https://substack.thewebscraping.club/p/scrapling-hands-on-guide?utm_source=github&utm_medium=repo&utm_campaign=scrapling" target="_blank">The Web Scraping Club で Scrapling の詳細レビュー</a>(2025年11月)をお読みください。Web スクレイピング専門の No.1 ニュースレターです。
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td width="200">
|
||||
<a href="https://proxy-seller.com/?partner=CU9CAA5TBYFFT2" target="_blank" title="Proxy-Seller provides reliable proxy infrastructure for Web Scraping">
|
||||
<img src="https://raw.githubusercontent.com/D4Vinci/Scrapling/main/images/ProxySeller.png">
|
||||
</a>
|
||||
</td>
|
||||
<td>
|
||||
<a href="https://proxy-seller.com/?partner=CU9CAA5TBYFFT2" target="_blank">Proxy-Seller</a> は Web スクレイピング向けの信頼性の高いプロキシインフラを提供しています。IPv4、IPv6、ISP、レジデンシャル、モバイルプロキシに対応し、安定したパフォーマンス、幅広い地理的カバレッジ、企業規模のデータ収集に柔軟なプランを備えています。
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
<i><sub>ここに広告を表示したいですか?[こちら](https://github.com/sponsors/D4Vinci/sponsorships?tier_id=586646)をクリック</sub></i>
|
||||
|
||||
+14
-4
@@ -34,17 +34,17 @@
|
||||
</p>
|
||||
|
||||
<p align="center">
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/parsing/selection/"><strong>선택 메서드</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/parsing/selection.html"><strong>선택 메서드</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/fetching/choosing/"><strong>Fetcher 선택 가이드</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/fetching/choosing.html"><strong>Fetcher 선택 가이드</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/spiders/architecture.html"><strong>Spider</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/spiders/proxy-blocking.html"><strong>프록시 로테이션</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/cli/overview/"><strong>CLI</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/cli/overview.html"><strong>CLI</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/ai/mcp-server/"><strong>MCP 서버</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/ai/mcp-server.html"><strong>MCP 서버</strong></a>
|
||||
</p>
|
||||
|
||||
Scrapling은 단일 요청부터 대규모 크롤링까지 모든 것을 처리하는 적응형 Web Scraping 프레임워크입니다.
|
||||
@@ -159,6 +159,16 @@ MySpider().start()
|
||||
<a href="https://substack.thewebscraping.club/p/scrapling-hands-on-guide?utm_source=github&utm_medium=repo&utm_campaign=scrapling" target="_blank">The Web Scraping Club에서 Scrapling의 전체 리뷰</a>(2025년 11월)를 읽어보세요. 웹 스크래핑 전문 No.1 뉴스레터입니다.
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td width="200">
|
||||
<a href="https://proxy-seller.com/?partner=CU9CAA5TBYFFT2" target="_blank" title="Proxy-Seller provides reliable proxy infrastructure for Web Scraping">
|
||||
<img src="https://raw.githubusercontent.com/D4Vinci/Scrapling/main/images/ProxySeller.png">
|
||||
</a>
|
||||
</td>
|
||||
<td>
|
||||
<a href="https://proxy-seller.com/?partner=CU9CAA5TBYFFT2" target="_blank">Proxy-Seller</a>는 웹 스크래핑을 위한 안정적인 프록시 인프라를 제공합니다. IPv4, IPv6, ISP, 주거용 및 모바일 프록시를 지원하며, 안정적인 성능, 광범위한 지역 커버리지, 기업 규모의 데이터 수집을 위한 유연한 요금제를 갖추고 있습니다.
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
<i><sub>여기에 광고를 게재하고 싶으신가요? [여기](https://github.com/sponsors/D4Vinci/sponsorships?tier_id=586646)를 클릭하세요</sub></i>
|
||||
|
||||
+14
-4
@@ -34,17 +34,17 @@
|
||||
</p>
|
||||
|
||||
<p align="center">
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/parsing/selection/"><strong>Методы выбора</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/parsing/selection.html"><strong>Методы выбора</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/fetching/choosing/"><strong>Выбор Fetcher</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/fetching/choosing.html"><strong>Выбор Fetcher</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/spiders/architecture.html"><strong>Пауки</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/spiders/proxy-blocking.html"><strong>Ротация прокси</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/cli/overview/"><strong>CLI</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/cli/overview.html"><strong>CLI</strong></a>
|
||||
·
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/ai/mcp-server/"><strong>Режим MCP</strong></a>
|
||||
<a href="https://scrapling.readthedocs.io/en/latest/ai/mcp-server.html"><strong>Режим MCP</strong></a>
|
||||
</p>
|
||||
|
||||
Scrapling — это адаптивный фреймворк для Web Scraping, который берёт на себя всё: от одного запроса до полномасштабного обхода сайтов.
|
||||
@@ -162,6 +162,16 @@ MySpider().start()
|
||||
Прочитайте полный обзор <a href="https://substack.thewebscraping.club/p/scrapling-hands-on-guide?utm_source=github&utm_medium=repo&utm_campaign=scrapling" target="_blank">Scrapling на The Web Scraping Club</a> (ноябрь 2025) — рассылка №1, посвящённая веб-скрейпингу.
|
||||
</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td width="200">
|
||||
<a href="https://proxy-seller.com/?partner=CU9CAA5TBYFFT2" target="_blank" title="Proxy-Seller provides reliable proxy infrastructure for Web Scraping">
|
||||
<img src="https://raw.githubusercontent.com/D4Vinci/Scrapling/main/images/ProxySeller.png">
|
||||
</a>
|
||||
</td>
|
||||
<td>
|
||||
<a href="https://proxy-seller.com/?partner=CU9CAA5TBYFFT2" target="_blank">Proxy-Seller</a> предоставляет надёжную прокси-инфраструктуру для веб-скрейпинга: IPv4, IPv6, ISP, резидентные и мобильные прокси со стабильной производительностью, широким географическим покрытием и гибкими тарифами для сбора данных в масштабах бизнеса.
|
||||
</td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
<i><sub>Хотите показать здесь свою рекламу? Нажмите [здесь](https://github.com/sponsors/D4Vinci/sponsorships?tier_id=586646)</sub></i>
|
||||
|
||||
+33
-7
@@ -46,16 +46,42 @@ MySpider().start()
|
||||
|
||||
## Top Sponsors
|
||||
|
||||
<style>
|
||||
.ad {
|
||||
width:240px;
|
||||
height:100px;
|
||||
}
|
||||
|
||||
</style>
|
||||
|
||||
<!-- sponsors -->
|
||||
<div style="text-align: center;">
|
||||
<a href="https://hypersolutions.co/?utm_source=github&utm_medium=readme&utm_campaign=scrapling" target="_blank" title="Bot Protection Bypass API for Akamai, DataDome, Incapsula & Kasada"><img src="https://raw.githubusercontent.com/D4Vinci/Scrapling/main/images/HyperSolutions.png"></a>
|
||||
<a href="https://birdproxies.com/t/scrapling" target="_blank" title="At Bird Proxies, we eliminate your pains such as banned IPs, geo restriction, and high costs so you can focus on your work."><img src="https://raw.githubusercontent.com/D4Vinci/Scrapling/main/images/BirdProxies.jpg"></a>
|
||||
<a href="https://evomi.com?utm_source=github&utm_medium=banner&utm_campaign=d4vinci-scrapling" target="_blank" title="Evomi is your Swiss Quality Proxy Provider, starting at $0.49/GB"><img src="https://raw.githubusercontent.com/D4Vinci/Scrapling/main/images/evomi.png"></a>
|
||||
<a href="https://tikhub.io/?ref=KarimShoair" target="_blank" title="Unlock the Power of Social Media Data & AI"><img src="https://raw.githubusercontent.com/D4Vinci/Scrapling/main/images/TikHub.jpg"></a><a href="https://www.nsocks.com/?keyword=2p67aivg" target="_blank" title="Scalable Web Data Access for AI Applications"><img src="https://raw.githubusercontent.com/D4Vinci/Scrapling/main/images/nsocks.png"></a><a href="https://petrosky.io/d4vinci" target="_blank" title="PetroSky delivers cutting-edge VPS hosting."><img src="https://raw.githubusercontent.com/D4Vinci/Scrapling/main/images/petrosky.png"></a><a href="https://substack.thewebscraping.club/p/scrapling-hands-on-guide?utm_source=github&utm_medium=repo&utm_campaign=scrapling" target="_blank" title="The #1 newsletter dedicated to Web Scraping">
|
||||
<img src="https://raw.githubusercontent.com/D4Vinci/Scrapling/main/images/TWSC.png">
|
||||
<a href="https://hypersolutions.co/?utm_source=github&utm_medium=readme&utm_campaign=scrapling" target="_blank" title="Bot Protection Bypass API for Akamai, DataDome, Incapsula & Kasada">
|
||||
<img src="https://raw.githubusercontent.com/D4Vinci/Scrapling/main/images/HyperSolutions.png" class="ad">
|
||||
</a>
|
||||
<br/><br/>
|
||||
|
||||
<a href="https://birdproxies.com/t/scrapling" target="_blank" title="At Bird Proxies, we eliminate your pains such as banned IPs, geo restriction, and high costs so you can focus on your work.">
|
||||
<img src="https://raw.githubusercontent.com/D4Vinci/Scrapling/main/images/BirdProxies.jpg" class="ad">
|
||||
</a>
|
||||
<a href="https://evomi.com?utm_source=github&utm_medium=banner&utm_campaign=d4vinci-scrapling" target="_blank" title="Evomi is your Swiss Quality Proxy Provider, starting at $0.49/GB">
|
||||
<img src="https://raw.githubusercontent.com/D4Vinci/Scrapling/main/images/evomi.png" class="ad">
|
||||
</a>
|
||||
<a href="https://tikhub.io/?ref=KarimShoair" target="_blank" title="Unlock the Power of Social Media Data & AI">
|
||||
<img src="https://raw.githubusercontent.com/D4Vinci/Scrapling/main/images/TikHub.jpg" class="ad">
|
||||
</a>
|
||||
<a href="https://www.nsocks.com/?keyword=2p67aivg" target="_blank" title="Scalable Web Data Access for AI Applications">
|
||||
<img src="https://raw.githubusercontent.com/D4Vinci/Scrapling/main/images/nsocks.png" class="ad">
|
||||
</a>
|
||||
<a href="https://petrosky.io/d4vinci" target="_blank" title="PetroSky delivers cutting-edge VPS hosting.">
|
||||
<img src="https://raw.githubusercontent.com/D4Vinci/Scrapling/main/images/petrosky.png" class="ad">
|
||||
</a>
|
||||
<a href="https://substack.thewebscraping.club/p/scrapling-hands-on-guide?utm_source=github&utm_medium=repo&utm_campaign=scrapling" target="_blank" title="The #1 newsletter dedicated to Web Scraping">
|
||||
<img src="https://raw.githubusercontent.com/D4Vinci/Scrapling/main/images/TWSC.png" class="ad">
|
||||
</a>
|
||||
<a href="https://proxy-seller.com/?partner=CU9CAA5TBYFFT2" target="_blank" title="Proxy-Seller provides reliable proxy infrastructure for Web Scraping">
|
||||
<img src="https://raw.githubusercontent.com/D4Vinci/Scrapling/main/images/ProxySeller.png" class="ad">
|
||||
</a>
|
||||
<br />
|
||||
<br />
|
||||
</div>
|
||||
<!-- /sponsors -->
|
||||
|
||||
|
||||
@@ -1,4 +1,4 @@
|
||||
zensical>=0.0.24
|
||||
zensical>=0.0.27
|
||||
mkdocstrings>=1.0.3
|
||||
mkdocstrings-python>=2.0.3
|
||||
griffe-inherited-docstrings>=1.1.3
|
||||
|
||||
@@ -39,30 +39,30 @@ Unless money is irrelevant to you, you will try to find less expensive approache
|
||||
Scrapling can handle almost all issues you will face during Web Scraping, and the following updates will cover the rest carefully.
|
||||
|
||||
### Solving issue T1: Rapidly changing website structures
|
||||
That's why the [adaptive](https://scrapling.readthedocs.io/en/latest/parsing/adaptive/) feature was made. You knew I would talk about it, and here we are :)
|
||||
That's why the [adaptive](https://scrapling.readthedocs.io/en/latest/parsing/adaptive.html) feature was made. You knew I would talk about it, and here we are :)
|
||||
|
||||
While Web Scraping, if you have the `adaptive` feature enabled, you can save any element's unique properties so you can find it again later when the website's structure changes. The most frustrating thing about changes is that anything about an element can change, so there's nothing to rely on.
|
||||
|
||||
That's how the adaptive feature works: it stores everything unique about an element. When the website structure changes, it returns the element with the highest similarity score of the previous element.
|
||||
|
||||
I have already explained this in more detail, with many examples. Read more from [here](https://scrapling.readthedocs.io/en/latest/parsing/adaptive/#how-the-adaptive-feature-works).
|
||||
I have already explained this in more detail, with many examples. Read more from [here](https://scrapling.readthedocs.io/en/latest/parsing/adaptive.html#how-the-adaptive-feature-works).
|
||||
|
||||
### Solving issue T2: Unstable selectors
|
||||
If you have been doing Web scraping for a long enough time, you have likely experienced this once. I'm referring to a website that employs poor design patterns, built on raw HTML without any IDs/classes, or uses random class names with nothing else to rely on, etc...
|
||||
|
||||
In these cases, standard selection methods with CSS/XPath selectors won't be optimal, and that's why Scrapling provides three more methods for Selection:
|
||||
|
||||
1. [Selection by element content](https://scrapling.readthedocs.io/en/latest/parsing/selection/#text-content-selection): Through text content (`find_by_text`) or regex that matches text content (`find_by_regex`)
|
||||
2. [Selecting elements similar to another element](https://scrapling.readthedocs.io/en/latest/parsing/selection/#finding-similar-elements): You find an element, and we will do the rest!
|
||||
3. [Selecting elements by filters](https://scrapling.readthedocs.io/en/latest/parsing/selection/#filters-based-searching): You specify conditions/filters that this element must fulfill, we find it!
|
||||
1. [Selection by element content](https://scrapling.readthedocs.io/en/latest/parsing/selection.html#text-content-selection): Through text content (`find_by_text`) or regex that matches text content (`find_by_regex`)
|
||||
2. [Selecting elements similar to another element](https://scrapling.readthedocs.io/en/latest/parsing/selection.html#finding-similar-elements): You find an element, and we will do the rest!
|
||||
3. [Selecting elements by filters](https://scrapling.readthedocs.io/en/latest/parsing/selection.html#filters-based-searching): You specify conditions/filters that this element must fulfill, we find it!
|
||||
|
||||
There is no need to explain any of these; click on the links, and it will be clear how Scrapling solves this.
|
||||
|
||||
### Solving issue T3: Increasingly complex anti-bot measures
|
||||
It's well known that creating an undetectable spider requires more than residential/mobile proxies and human-like behavior. It also needs a hard-to-detect browser, which Scrapling provides two main options to solve:
|
||||
|
||||
1. [DynamicFetcher](https://scrapling.readthedocs.io/en/latest/fetching/dynamic/) — This fetcher provides flexible browser automation with multiple configuration options and little under-the-hood stealth improvements.
|
||||
2. [StealthyFetcher](https://scrapling.readthedocs.io/en/latest/fetching/stealthy/) — Because we live in a harsh world and you need to take [full measure instead of half-measures](https://www.youtube.com/watch?v=7BE4QcwX4dU), `StealthyFetcher` was born. This fetcher uses our stealthy browser -- a version of [DynamicFetcher](https://scrapling.readthedocs.io/en/latest/fetching/dynamic/) that nearly bypasses all annoying anti-protections, provides tools to handle the rest, and automatically bypasses all types of Cloudflare's Turnstile/Interstitial!
|
||||
1. [DynamicFetcher](https://scrapling.readthedocs.io/en/latest/fetching/dynamic.html) — This fetcher provides flexible browser automation with multiple configuration options and little under-the-hood stealth improvements.
|
||||
2. [StealthyFetcher](https://scrapling.readthedocs.io/en/latest/fetching/stealthy.html) — Because we live in a harsh world and you need to take [full measure instead of half-measures](https://www.youtube.com/watch?v=7BE4QcwX4dU), `StealthyFetcher` was born. This fetcher uses our stealthy browser -- a version of [DynamicFetcher](https://scrapling.readthedocs.io/en/latest/fetching/dynamic.html) that nearly bypasses all annoying anti-protections, provides tools to handle the rest, and automatically bypasses all types of Cloudflare's Turnstile/Interstitial!
|
||||
|
||||
We keep improving these two with each update, so stay tuned :)
|
||||
|
||||
|
||||
Binary file not shown.
|
After Width: | Height: | Size: 26 KiB |
@@ -112,10 +112,10 @@ class TextHandler(str):
|
||||
def get(self, default=None): # pragma: no cover
|
||||
return self
|
||||
|
||||
def get_all(self): # pragma: no cover
|
||||
def getall(self): # pragma: no cover
|
||||
return self
|
||||
|
||||
extract = get_all
|
||||
extract = getall
|
||||
extract_first = get
|
||||
|
||||
def json(self) -> Dict:
|
||||
@@ -279,7 +279,7 @@ class TextHandlers(List[TextHandler]):
|
||||
return self
|
||||
|
||||
extract_first = get
|
||||
get_all = extract
|
||||
getall = extract
|
||||
|
||||
|
||||
class AttributesHandler(Mapping[str, _TextHandlerType]):
|
||||
|
||||
@@ -187,34 +187,34 @@ class ResponseFactory:
|
||||
return history
|
||||
|
||||
@classmethod
|
||||
def _get_page_content(cls, page: SyncPage) -> str:
|
||||
def _get_page_content(cls, page: SyncPage, max_retries: int = 20) -> str:
|
||||
"""
|
||||
A workaround for the Playwright issue with `page.content()` on Windows. Ref.: https://github.com/microsoft/playwright/issues/16108
|
||||
:param page: The page to extract content from.
|
||||
:param max_retries: Maximum number of retry attempts before raising `RuntimeError`.
|
||||
:return:
|
||||
"""
|
||||
while True:
|
||||
for _ in range(max_retries):
|
||||
try:
|
||||
return page.content() or ""
|
||||
except PlaywrightError:
|
||||
page.wait_for_timeout(500)
|
||||
continue
|
||||
return "" # pyright: ignore
|
||||
raise RuntimeError(f"Failed to retrieve the page content after retrying for {max_retries * 500}ms.")
|
||||
|
||||
@classmethod
|
||||
async def _get_async_page_content(cls, page: AsyncPage) -> str:
|
||||
async def _get_async_page_content(cls, page: AsyncPage, max_retries: int = 20) -> str:
|
||||
"""
|
||||
A workaround for the Playwright issue with `page.content()` on Windows. Ref.: https://github.com/microsoft/playwright/issues/16108
|
||||
:param page: The page to extract content from.
|
||||
:param max_retries: Maximum number of retry attempts before raising `RuntimeError`.
|
||||
:return:
|
||||
"""
|
||||
while True:
|
||||
for _ in range(max_retries):
|
||||
try:
|
||||
return (await page.content()) or ""
|
||||
except PlaywrightError:
|
||||
await page.wait_for_timeout(500)
|
||||
continue
|
||||
return "" # pyright: ignore
|
||||
raise RuntimeError(f"Failed to retrieve the page content after retrying for {max_retries * 500}ms.")
|
||||
|
||||
@classmethod
|
||||
async def from_async_playwright_response(
|
||||
|
||||
@@ -205,7 +205,7 @@ class CrawlerEngine:
|
||||
Returns True if successfully restored, False otherwise.
|
||||
"""
|
||||
if not self._checkpoint_system_enabled:
|
||||
raise
|
||||
return False
|
||||
|
||||
data = await self._checkpoint_manager.load()
|
||||
if data is None:
|
||||
|
||||
@@ -112,10 +112,12 @@ class SessionManager:
|
||||
client = session._client
|
||||
|
||||
if isinstance(client, _ASyncSessionLogic):
|
||||
kwargs = request._session_kwargs.copy()
|
||||
method = cast(SUPPORTED_HTTP_METHODS, kwargs.pop("method", "GET"))
|
||||
response = await client._make_request(
|
||||
method=cast(SUPPORTED_HTTP_METHODS, request._session_kwargs.pop("method", "GET")),
|
||||
method=method,
|
||||
url=request.url,
|
||||
**request._session_kwargs,
|
||||
**kwargs,
|
||||
)
|
||||
else:
|
||||
# Sync session or other types - shouldn't happen in async context
|
||||
|
||||
@@ -250,6 +250,68 @@ class TestTextHandlerAdvanced:
|
||||
matches = text3.re(r"He l lo", clean_match=True, case_sensitive=False)
|
||||
assert len(matches) == 1
|
||||
|
||||
def test_text_handler_regex_check_match(self):
|
||||
"""Test TextHandler.re() with check_match=True returns bool"""
|
||||
text = TextHandler("Price: $10.99")
|
||||
assert text.re(r"\$[\d.]+", check_match=True) is True
|
||||
assert text.re(r"no-match-pattern", check_match=True) is False
|
||||
|
||||
def test_text_handler_regex_replace_entities_false(self):
|
||||
"""Test TextHandler.re() with replace_entities=False preserves entities"""
|
||||
text = TextHandler("Hello & World")
|
||||
results = text.re(r"&", replace_entities=False)
|
||||
assert len(results) == 1
|
||||
assert results[0] == "&"
|
||||
|
||||
def test_text_handler_regex_with_groups(self):
|
||||
"""Test TextHandler.re() with capture groups flattens results"""
|
||||
text = TextHandler("name=Alice age=30 name=Bob age=25")
|
||||
results = text.re(r"name=(\w+) age=(\d+)")
|
||||
assert len(results) == 4
|
||||
assert "Alice" in results
|
||||
assert "30" in results
|
||||
|
||||
def test_text_handler_re_first_with_default(self):
|
||||
"""Test TextHandler.re_first() returns default when no match"""
|
||||
text = TextHandler("no numbers here")
|
||||
result = text.re_first(r"\d+", default="N/A")
|
||||
assert result == "N/A"
|
||||
|
||||
def test_text_handler_re_first_returns_first_match(self):
|
||||
"""Test TextHandler.re_first() returns first match"""
|
||||
text = TextHandler("a1 b2 c3")
|
||||
result = text.re_first(r"\d")
|
||||
assert result == "1"
|
||||
assert isinstance(result, TextHandler)
|
||||
|
||||
def test_text_handler_clean_with_entities(self):
|
||||
"""Test TextHandler.clean() with remove_entities=True"""
|
||||
text = TextHandler("Hello\t&\nWorld")
|
||||
cleaned = text.clean(remove_entities=True)
|
||||
assert "&" not in cleaned
|
||||
assert "&" in cleaned
|
||||
assert "\t" not in cleaned
|
||||
assert "\n" not in cleaned
|
||||
|
||||
def test_text_handler_clean_without_entities(self):
|
||||
"""Test TextHandler.clean() preserves entities by default"""
|
||||
text = TextHandler("Hello\t&\nWorld")
|
||||
cleaned = text.clean(remove_entities=False)
|
||||
assert "&" in cleaned
|
||||
|
||||
def test_text_handler_json_valid(self):
|
||||
"""Test TextHandler.json() with valid JSON"""
|
||||
text = TextHandler('{"key": "value", "num": 42}')
|
||||
data = text.json()
|
||||
assert data["key"] == "value"
|
||||
assert data["num"] == 42
|
||||
|
||||
def test_text_handler_json_invalid(self):
|
||||
"""Test TextHandler.json() raises on invalid JSON"""
|
||||
text = TextHandler("not json")
|
||||
with pytest.raises(Exception):
|
||||
text.json()
|
||||
|
||||
def test_text_handlers_operations(self):
|
||||
"""Test TextHandlers list operations"""
|
||||
handlers = TextHandlers([
|
||||
@@ -266,6 +328,37 @@ class TestTextHandlerAdvanced:
|
||||
assert handlers.get("default") == "First"
|
||||
assert TextHandlers([]).get("default") == "default"
|
||||
|
||||
def test_text_handlers_re(self):
|
||||
"""Test TextHandlers.re() flattens results across all elements"""
|
||||
handlers = TextHandlers([
|
||||
TextHandler("a1 b2"),
|
||||
TextHandler("c3 d4"),
|
||||
])
|
||||
results = handlers.re(r"[a-z]\d")
|
||||
assert isinstance(results, TextHandlers)
|
||||
assert len(results) == 4
|
||||
assert results[0] == "a1"
|
||||
assert results[3] == "d4"
|
||||
|
||||
def test_text_handlers_re_empty(self):
|
||||
"""Test TextHandlers.re() on empty list"""
|
||||
handlers = TextHandlers([])
|
||||
results = handlers.re(r"\d+")
|
||||
assert isinstance(results, TextHandlers)
|
||||
assert len(results) == 0
|
||||
|
||||
def test_text_handlers_re_no_matches(self):
|
||||
"""Test TextHandlers.re() when no element matches"""
|
||||
handlers = TextHandlers([TextHandler("abc"), TextHandler("def")])
|
||||
results = handlers.re(r"\d+")
|
||||
assert len(results) == 0
|
||||
|
||||
def test_text_handlers_extract(self):
|
||||
"""Test TextHandlers.extract() returns self"""
|
||||
handlers = TextHandlers([TextHandler("a"), TextHandler("b")])
|
||||
assert handlers.extract() is handlers
|
||||
assert handlers.getall() is handlers
|
||||
|
||||
|
||||
class TestSelectorsAdvanced:
|
||||
"""Test advanced Selectors functionality"""
|
||||
|
||||
@@ -613,8 +613,7 @@ class TestCheckpointMethods:
|
||||
@pytest.mark.asyncio
|
||||
async def test_restore_from_checkpoint_raises_when_disabled(self):
|
||||
engine = _make_engine() # no crawldir → checkpoint disabled
|
||||
with pytest.raises(RuntimeError):
|
||||
await engine._restore_from_checkpoint()
|
||||
assert (await engine._restore_from_checkpoint()) is False
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
@@ -1,9 +1,12 @@
|
||||
"""Tests for the SessionManager class."""
|
||||
|
||||
from unittest.mock import AsyncMock, PropertyMock
|
||||
|
||||
from scrapling.core._types import Any
|
||||
import pytest
|
||||
|
||||
from scrapling.spiders.session import SessionManager
|
||||
from scrapling.spiders.request import Request
|
||||
|
||||
|
||||
class MockSession: # type: ignore[type-arg]
|
||||
@@ -350,3 +353,57 @@ class TestSessionManagerIntegration:
|
||||
# After close - all inactive
|
||||
await manager.close()
|
||||
assert all(not s._is_alive for s in sessions)
|
||||
|
||||
|
||||
class TestSessionManagerFetch:
|
||||
"""Test SessionManager fetch behavior."""
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_fetch_preserves_request_method(self):
|
||||
"""Test that fetch does not mutate request._session_kwargs.
|
||||
|
||||
Previously, fetch() used pop("method") which removed the method
|
||||
key from the original request dict. This caused retried requests
|
||||
(via request.copy()) to lose their HTTP method and fall back to GET.
|
||||
"""
|
||||
from scrapling.engines.static import _ASyncSessionLogic
|
||||
from scrapling.fetchers import FetcherSession
|
||||
from scrapling.engines.toolbelt.custom import Response
|
||||
|
||||
mock_response = Response(
|
||||
url="https://example.com",
|
||||
content=b"ok",
|
||||
status=200,
|
||||
reason="OK",
|
||||
cookies={},
|
||||
headers={"content-type": "text/html"},
|
||||
request_headers={},
|
||||
)
|
||||
mock_response.meta = {}
|
||||
|
||||
mock_client = AsyncMock(spec=_ASyncSessionLogic)
|
||||
mock_client._make_request = AsyncMock(return_value=mock_response)
|
||||
|
||||
mock_session = AsyncMock(spec=FetcherSession)
|
||||
mock_session._client = mock_client
|
||||
mock_session._is_alive = True
|
||||
|
||||
manager = SessionManager()
|
||||
manager._sessions["default"] = mock_session
|
||||
manager._default_session_id = "default"
|
||||
manager._started = True
|
||||
|
||||
request = Request("https://example.com", method="POST", data={"key": "value"})
|
||||
|
||||
assert request._session_kwargs["method"] == "POST"
|
||||
|
||||
await manager.fetch(request)
|
||||
|
||||
# method must still be present after fetch
|
||||
assert "method" in request._session_kwargs
|
||||
assert request._session_kwargs["method"] == "POST"
|
||||
|
||||
# verify the correct method was passed to _make_request
|
||||
mock_client._make_request.assert_called_once()
|
||||
call_kwargs = mock_client._make_request.call_args
|
||||
assert call_kwargs.kwargs["method"] == "POST"
|
||||
|
||||
Reference in New Issue
Block a user