diff --git a/.github/workflows/tests.yml b/.github/workflows/tests.yml
index f82605d..930d0a0 100644
--- a/.github/workflows/tests.yml
+++ b/.github/workflows/tests.yml
@@ -73,7 +73,7 @@ jobs:
- name: Install all browsers dependencies
run: |
python3 -m pip install --upgrade pip
- python3 -m pip install playwright==1.59.0 patchright==1.59.1
+ python3 -m pip install playwright==1.60.0 patchright==1.60.1
- name: Get Playwright version
id: playwright-version
diff --git a/README.md b/README.md
index 6c087b8..4507844 100644
--- a/README.md
+++ b/README.md
@@ -141,16 +141,6 @@ MySpider().start()
TikHub.io provides 900+ stable APIs across 16+ platforms including TikTok, X, YouTube & Instagram, with 40M+ datasets.
Also offers DISCOUNTED AI models - Claude, GPT, GEMINI & more up to 71% off.
-
-
-
-
-
- |
-
- Nsocks provides fast Residential and ISP proxies for developers and scrapers. Global IP coverage, high anonymity, smart rotation, and reliable performance for automation and data extraction. Use Xcrawl to simplify large-scale web crawling.
- |
-
|
@@ -172,16 +162,6 @@ MySpider().start()
Read a full review of Scrapling on The Web Scraping Club (Nov 2025), the #1 newsletter dedicated to Web Scraping.
|
-
-
-
-
-
- |
-
- Stable proxies for scraping, automation, and multi-accounting. Clean IPs, fast response, and reliable performance under load. Built for scalable workflows.
- |
-
|
@@ -192,22 +172,38 @@ MySpider().start()
Swiftproxy provides scalable residential proxies with 80M+ IPs across 195+ countries, delivering fast, reliable connections, automatic rotation, and strong anti-block performance. Free trial available.
|
+
+
+
+
+
+ |
+
+ 9Proxy provides residential proxies from just $0.018/IP or $0.68/GB. 20M+ IPs across 90+ countries. Sticky or rotating sessions, managed from desktop or mobile app.
+ |
+
+
+
+
+
+
+ |
+
+ NodeMaven - reliable proxy provider with the highest quality IP on the market. Use promo code SCRAPLING35 for 35% discount on proxies.
+ |
+
Do you want to show your ad here? Click [here](https://github.com/sponsors/D4Vinci/sponsorships?tier_id=586646)
# Sponsors
-
-
-
-
diff --git a/agent-skill/Scrapling-Skill.zip b/agent-skill/Scrapling-Skill.zip
index 22fbb32..1a8d15b 100644
Binary files a/agent-skill/Scrapling-Skill.zip and b/agent-skill/Scrapling-Skill.zip differ
diff --git a/agent-skill/Scrapling-Skill/SKILL.md b/agent-skill/Scrapling-Skill/SKILL.md
index 27c31c0..ab57268 100644
--- a/agent-skill/Scrapling-Skill/SKILL.md
+++ b/agent-skill/Scrapling-Skill/SKILL.md
@@ -1,7 +1,7 @@
---
name: scrapling-official
description: Scrape web pages using Scrapling with anti-bot bypass (like Cloudflare Turnstile), stealth headless browsing, spiders framework, adaptive scraping, and JavaScript rendering. Use when asked to scrape, crawl, or extract data from websites; web_fetch fails; the site has anti-bot protections; write Python code to scrape/crawl; or write spiders.
-version: "0.4.8"
+version: "0.4.9"
license: Complete terms in LICENSE.txt
metadata:
homepage: "https://scrapling.readthedocs.io/en/latest/index.html"
@@ -40,7 +40,7 @@ Blazing fast crawls with real-time stats and streaming. Built by Web Scrapers fo
Create a virtual Python environment through any way available, like `venv`, then inside the environment do:
-`pip install "scrapling[all]>=0.4.8"`
+`pip install "scrapling[all]>=0.4.9"`
Then do this to download all the browsers' dependencies:
diff --git a/agent-skill/Scrapling-Skill/examples/README.md b/agent-skill/Scrapling-Skill/examples/README.md
index d0f9a2b..85de486 100644
--- a/agent-skill/Scrapling-Skill/examples/README.md
+++ b/agent-skill/Scrapling-Skill/examples/README.md
@@ -9,7 +9,7 @@ All examples collect **all 100 quotes across 10 pages**.
Make sure Scrapling is installed:
```bash
-pip install "scrapling[all]>=0.4.8"
+pip install "scrapling[all]>=0.4.9"
scrapling install --force
```
diff --git a/agent-skill/Scrapling-Skill/references/mcp-server.md b/agent-skill/Scrapling-Skill/references/mcp-server.md
index 48a2d10..6aaa014 100644
--- a/agent-skill/Scrapling-Skill/references/mcp-server.md
+++ b/agent-skill/Scrapling-Skill/references/mcp-server.md
@@ -208,7 +208,7 @@ Docker alternative:
```bash
docker pull pyd4vinci/scrapling
-docker run -i --rm scrapling mcp
+docker run -i --rm pyd4vinci/scrapling mcp
```
The MCP server name when registering with a client is `ScraplingServer`. The command is the path to the `scrapling` binary and the argument is `mcp`.
\ No newline at end of file
diff --git a/docs/README_AR.md b/docs/README_AR.md
index 7c28c79..2070bc9 100644
--- a/docs/README_AR.md
+++ b/docs/README_AR.md
@@ -137,16 +137,6 @@ MySpider().start()
TikHub.io يوفر أكثر من 900 واجهة API مستقرة عبر أكثر من 16 منصة تشمل TikTok و X و YouTube و Instagram، مع أكثر من 40 مليون مجموعة بيانات.
يقدم أيضاً نماذج ذكاء اصطناعي بأسعار مخفضة - Claude و GPT و GEMINI والمزيد بخصم يصل إلى 71%.
-
-
-
-
-
- |
-
- Nsocks يوفر بروكسيات سكنية و ISP سريعة للمطورين والسكرابرز. تغطية IP عالمية، إخفاء هوية عالي، تدوير ذكي، وأداء موثوق للأتمتة واستخراج البيانات. استخدم Xcrawl لتبسيط زحف الويب على نطاق واسع.
- |
-
|
@@ -168,16 +158,6 @@ MySpider().start()
اقرأ مراجعة كاملة عن Scrapling على The Web Scraping Club (نوفمبر 2025)، النشرة الإخبارية الأولى المخصصة لكشط الويب.
|
-
-
-
-
-
- |
-
- بروكسيات مستقرة للكشط والأتمتة وإدارة الحسابات المتعددة. عناوين IP نظيفة، استجابة سريعة، وأداء موثوق تحت الضغط. مصممة لسير العمل القابل للتوسع.
- |
-
|
@@ -188,23 +168,38 @@ MySpider().start()
يوفر Swiftproxy بروكسيات سكنية قابلة للتوسع مع أكثر من 80 مليون عنوان IP في أكثر من 195 دولة، ويقدم اتصالات سريعة وموثوقة، وتدوير تلقائي، وأداء قوي ضد الحظر. تجربة مجانية متاحة.
|
+
+
+
+
+
+ |
+
+ يوفر 9Proxy بروكسيات سكنية بدءًا من 0.018 دولار فقط لكل IP أو 0.68 دولار لكل جيجابايت. أكثر من 20 مليون عنوان IP في أكثر من 90 دولة. جلسات ثابتة أو متناوبة، تتم إدارتها من تطبيق سطح المكتب أو الجوال.
+ |
+
+
+
+
+
+
+ |
+
+ NodeMaven - مزود بروكسيات موثوق يقدم أعلى جودة IP في السوق. استخدم كود الخصم SCRAPLING35 للحصول على خصم 35% على البروكسيات.
+ |
+
هل تريد عرض إعلانك هنا؟ انقر [هنا](https://github.com/sponsors/D4Vinci/sponsorships?tier_id=586646)
# الرعاة
-
-
-
-
-
diff --git a/docs/README_CN.md b/docs/README_CN.md
index b4a1dd1..0750bcc 100644
--- a/docs/README_CN.md
+++ b/docs/README_CN.md
@@ -137,16 +137,6 @@ MySpider().start()
TikHub.io 提供覆盖 16+ 平台(包括 TikTok、X、YouTube 和 Instagram)的 900+ 稳定 API,拥有 4000 万+ 数据集。
还提供优惠 AI 模型 - Claude、GPT、GEMINI 等,最高优惠 71%。
-
-
-
-
-
- |
-
- Nsocks 提供面向开发者和爬虫的快速住宅和 ISP 代理。全球 IP 覆盖、高匿名性、智能轮换,以及可靠的自动化和数据提取性能。使用 Xcrawl 简化大规模网页爬取。
- |
-
|
@@ -168,16 +158,6 @@ MySpider().start()
阅读 The Web Scraping Club 上关于 Scrapling 的完整评测(2025 年 11 月),这是排名第一的网页抓取专业通讯。
|
-
-
-
-
-
- |
-
- 稳定的代理,适用于数据抓取、自动化和多账号管理。干净的 IP、快速响应、高负载下可靠的性能。专为可扩展的工作流程而构建。
- |
-
|
@@ -188,22 +168,38 @@ MySpider().start()
Swiftproxy 提供可扩展的住宅代理,覆盖 195+ 国家/地区的 8000 万+ IP,提供快速可靠的连接、自动轮换和强大的反屏蔽性能。提供免费试用。
|
+
+
+
+
+
+ |
+
+ 9Proxy 提供住宅代理,价格低至每个 IP 仅 $0.018 或每 GB $0.68。覆盖 90+ 国家/地区的 2000 万+ IP。支持固定或轮换会话,可通过桌面或移动应用进行管理。
+ |
+
+
+
+
+
+
+ |
+
+ NodeMaven - 市场上 IP 质量最高的可靠代理提供商。使用优惠码 SCRAPLING35 可享代理 35% 折扣。
+ |
+
想在这里展示您的广告吗?点击 [这里](https://github.com/sponsors/D4Vinci/sponsorships?tier_id=586646)
# 赞助商
-
-
-
-
diff --git a/docs/README_DE.md b/docs/README_DE.md
index ee133de..ac6e533 100644
--- a/docs/README_DE.md
+++ b/docs/README_DE.md
@@ -137,16 +137,6 @@ MySpider().start()
TikHub.io bietet über 900 stabile APIs auf mehr als 16 Plattformen, darunter TikTok, X, YouTube und Instagram, mit über 40 Mio. Datensätzen.
Bietet außerdem vergünstigte KI-Modelle - Claude, GPT, GEMINI und mehr mit bis zu 71% Rabatt.
-
-
-
-
-
- |
-
- Nsocks bietet schnelle Residential- und ISP-Proxies für Entwickler und Scraper. Globale IP-Abdeckung, hohe Anonymität, intelligente Rotation und zuverlässige Leistung für Automatisierung und Datenextraktion. Verwenden Sie Xcrawl, um großflächiges Web-Crawling zu vereinfachen.
- |
-
|
@@ -168,16 +158,6 @@ MySpider().start()
Lesen Sie eine vollständige Rezension von Scrapling auf The Web Scraping Club (Nov. 2025), dem führenden Newsletter für Web Scraping.
|
-
-
-
-
-
- |
-
- Stabile Proxys für Scraping, Automatisierung und Multi-Accounting. Saubere IPs, schnelle Reaktionszeiten und zuverlässige Leistung unter Last. Entwickelt für skalierbare Workflows.
- |
-
|
@@ -188,22 +168,38 @@ MySpider().start()
Swiftproxy bietet skalierbare Residential-Proxys mit über 80 Mio. IPs in mehr als 195 Ländern und liefert schnelle, zuverlässige Verbindungen, automatische Rotation und starke Anti-Block-Leistung. Kostenlose Testversion verfügbar.
|
+
+
+
+
+
+ |
+
+ 9Proxy bietet Residential-Proxys ab nur 0,018 $/IP oder 0,68 $/GB. Über 20 Mio. IPs in mehr als 90 Ländern. Sticky oder rotierende Sessions, verwaltet über die Desktop- oder Mobile-App.
+ |
+
+
+
+
+
+
+ |
+
+ NodeMaven - zuverlässiger Proxy-Anbieter mit der höchsten IP-Qualität auf dem Markt. Nutze den Promo-Code SCRAPLING35 für 35% Rabatt auf Proxys.
+ |
+
Möchten Sie Ihre Anzeige hier zeigen? Klicken Sie [hier](https://github.com/sponsors/D4Vinci/sponsorships?tier_id=586646)
# Sponsoren
-
-
-
-
diff --git a/docs/README_ES.md b/docs/README_ES.md
index 056d7ba..1af1f9a 100644
--- a/docs/README_ES.md
+++ b/docs/README_ES.md
@@ -137,16 +137,6 @@ MySpider().start()
TikHub.io ofrece más de 900 APIs estables en más de 16 plataformas, incluyendo TikTok, X, YouTube e Instagram, con más de 40M de conjuntos de datos.
También ofrece modelos de IA con descuento - Claude, GPT, GEMINI y más con hasta un 71% de descuento.
-
-
-
-
-
- |
-
- Nsocks ofrece proxies residenciales e ISP rápidos para desarrolladores y scrapers. Cobertura IP global, alto anonimato, rotación inteligente y rendimiento fiable para automatización y extracción de datos. Usa Xcrawl para simplificar el crawling web a gran escala.
- |
-
|
@@ -168,16 +158,6 @@ MySpider().start()
Lee una reseña completa de Scrapling en The Web Scraping Club (nov. 2025), el boletín número uno dedicado al Web Scraping.
|
-
-
-
-
-
- |
-
- Proxies estables para scraping, automatización y multicuentas. IPs limpias, respuesta rápida y rendimiento fiable bajo carga. Diseñado para flujos de trabajo escalables.
- |
-
|
@@ -188,22 +168,38 @@ MySpider().start()
Swiftproxy ofrece proxies residenciales escalables con más de 80 millones de IPs en más de 195 países, brindando conexiones rápidas y fiables, rotación automática y un sólido rendimiento anti-bloqueo. Prueba gratuita disponible.
|
+
+
+
+
+
+ |
+
+ 9Proxy ofrece proxies residenciales desde solo $0,018/IP o $0,68/GB. Más de 20 millones de IPs en más de 90 países. Sesiones fijas o rotativas, gestionadas desde la aplicación de escritorio o móvil.
+ |
+
+
+
+
+
+
+ |
+
+ NodeMaven - proveedor de proxies confiable con la mayor calidad de IP del mercado. Usa el código promocional SCRAPLING35 para obtener un 35% de descuento en proxies.
+ |
+
¿Quieres mostrar tu anuncio aquí? Haz clic [aquí](https://github.com/sponsors/D4Vinci/sponsorships?tier_id=586646)
# Patrocinadores
-
-
-
-
diff --git a/docs/README_FR.md b/docs/README_FR.md
index 862e989..c6d76ed 100644
--- a/docs/README_FR.md
+++ b/docs/README_FR.md
@@ -137,16 +137,6 @@ MySpider().start()
TikHub.io propose plus de 900 APIs stables sur plus de 16 plateformes, dont TikTok, X, YouTube et Instagram, avec plus de 40M de jeux de données.
Propose également des modèles IA à prix réduit - Claude, GPT, GEMINI et plus, jusqu'à 71% de réduction.
-
-
-
-
-
- |
-
- Nsocks fournit des proxies résidentiels et ISP rapides pour les développeurs et les scrapeurs. Couverture IP mondiale, anonymat élevé, rotation intelligente et performances fiables pour l'automatisation et l'extraction de données. Utilisez Xcrawl pour simplifier le crawling web à grande échelle.
- |
-
|
@@ -168,16 +158,6 @@ MySpider().start()
Lisez une critique complète de Scrapling sur The Web Scraping Club (nov. 2025), la newsletter n°1 dédiée au Web Scraping.
|
-
-
-
-
-
- |
-
- Des proxys stables pour le scraping, l'automatisation et la gestion multi-comptes. Des IPs propres, une réponse rapide et des performances fiables sous charge. Conçu pour des flux de travail évolutifs.
- |
-
|
@@ -188,22 +168,38 @@ MySpider().start()
Swiftproxy propose des proxys résidentiels évolutifs avec plus de 80 millions d'IPs dans plus de 195 pays, offrant des connexions rapides et fiables, une rotation automatique et de solides performances anti-blocage. Essai gratuit disponible.
|
+
+
+
+
+
+ |
+
+ 9Proxy propose des proxys résidentiels à partir de seulement 0,018 $/IP ou 0,68 $/Go. Plus de 20 millions d'IPs dans plus de 90 pays. Sessions fixes ou rotatives, gérées depuis l'application de bureau ou mobile.
+ |
+
+
+
+
+
+
+ |
+
+ NodeMaven - fournisseur de proxys fiable offrant la meilleure qualité d'IP du marché. Utilisez le code promo SCRAPLING35 pour obtenir 35% de réduction sur les proxys.
+ |
+
Vous souhaitez afficher votre publicité ici ? Cliquez [ici](https://github.com/sponsors/D4Vinci/sponsorships?tier_id=586646)
# Sponsors
-
-
-
-
diff --git a/docs/README_JP.md b/docs/README_JP.md
index a82b916..dcf44fc 100644
--- a/docs/README_JP.md
+++ b/docs/README_JP.md
@@ -137,16 +137,6 @@ MySpider().start()
TikHub.io は TikTok、X、YouTube、Instagram を含む 16 以上のプラットフォームで 900 以上の安定した API を提供し、4,000 万以上のデータセットを保有。
さらに 割引 AI モデルも提供 - Claude、GPT、GEMINI など最大 71% オフ。
-
-
-
-
-
- |
-
- Nsocks は開発者やスクレイパー向けの高速なレジデンシャルおよび ISP プロキシを提供。グローバル IP カバレッジ、高い匿名性、スマートなローテーション、自動化とデータ抽出のための信頼性の高いパフォーマンス。Xcrawl で大規模ウェブクローリングを簡素化。
- |
-
|
@@ -168,16 +158,6 @@ MySpider().start()
The Web Scraping Club で Scrapling の詳細レビュー(2025年11月)をお読みください。Web スクレイピング専門の No.1 ニュースレターです。
|
-
-
-
-
-
- |
-
- 安定したプロキシ。スクレイピング、自動化、マルチアカウント管理に対応。クリーンな IP、高速レスポンス、高負荷時でも信頼性の高いパフォーマンス。スケーラブルなワークフロー向けに設計。
- |
-
|
@@ -188,22 +168,38 @@ MySpider().start()
Swiftproxy は195カ国以上、8,000万以上のIPを備えたスケーラブルな住宅用プロキシを提供し、高速で信頼性の高い接続、自動ローテーション、強力なブロック回避性能を実現します。無料トライアルあり。
|
+
+
+
+
+
+ |
+
+ 9Proxy はIPあたり $0.018 またはGBあたり $0.68 からの住宅用プロキシを提供します。90カ国以上で2,000万以上のIPを保有。固定セッションまたはローテーションセッションをデスクトップまたはモバイルアプリで管理できます。
+ |
+
+
+
+
+
+
+ |
+
+ NodeMaven - 市場最高品質のIPを提供する信頼性の高いプロキシプロバイダー。プロモコード SCRAPLING35 でプロキシが35%割引になります。
+ |
+
ここに広告を表示したいですか?[こちら](https://github.com/sponsors/D4Vinci/sponsorships?tier_id=586646)をクリック
# スポンサー
-
-
-
-
diff --git a/docs/README_KR.md b/docs/README_KR.md
index 9bf98c4..18d2e05 100644
--- a/docs/README_KR.md
+++ b/docs/README_KR.md
@@ -137,16 +137,6 @@ MySpider().start()
TikHub.io는 TikTok, X, YouTube, Instagram 등 16개 이상 플랫폼에서 900개 이상의 안정적인 API를 제공하며, 4,000만 이상의 데이터셋을 보유하고 있습니다.
할인된 AI 모델도 제공 - Claude, GPT, GEMINI 등 최대 71% 할인.
-
-
-
-
-
- |
-
- Nsocks는 개발자와 스크레이퍼를 위한 빠른 레지덴셜 및 ISP 프록시를 제공합니다. 글로벌 IP 커버리지, 높은 익명성, 스마트 로테이션, 자동화와 데이터 추출을 위한 안정적인 성능. Xcrawl로 대규모 웹 크롤링을 간소화하세요.
- |
-
|
@@ -168,16 +158,6 @@ MySpider().start()
The Web Scraping Club에서 Scrapling의 전체 리뷰(2025년 11월)를 읽어보세요. 웹 스크래핑 전문 No.1 뉴스레터입니다.
|
-
-
-
-
-
- |
-
- 안정적인 프록시. 스크래핑, 자동화, 멀티 계정 관리에 적합합니다. 깨끗한 IP, 빠른 응답, 높은 부하에서도 신뢰할 수 있는 성능. 확장 가능한 워크플로우를 위해 설계되었습니다.
- |
-
|
@@ -188,22 +168,38 @@ MySpider().start()
Swiftproxy는 195개국 이상에서 8천만 개 이상의 IP를 갖춘 확장 가능한 주거용 프록시를 제공하며, 빠르고 안정적인 연결, 자동 회전, 강력한 차단 방지 성능을 제공합니다. 무료 체험판 이용 가능.
|
+
+
+
+
+
+ |
+
+ 9Proxy는 IP당 $0.018 또는 GB당 $0.68의 저렴한 가격부터 시작하는 주거용 프록시를 제공합니다. 90개국 이상에서 2천만 개 이상의 IP를 보유하고 있으며, 고정 또는 회전 세션을 데스크톱이나 모바일 앱에서 관리할 수 있습니다.
+ |
+
+
+
+
+
+
+ |
+
+ NodeMaven - 시장에서 가장 높은 품질의 IP를 제공하는 신뢰할 수 있는 프록시 제공업체입니다. 프로모 코드 SCRAPLING35를 사용하면 프록시 35% 할인을 받을 수 있습니다.
+ |
+
여기에 광고를 게재하고 싶으신가요? [여기](https://github.com/sponsors/D4Vinci/sponsorships?tier_id=586646)를 클릭하세요
# 스폰서
-
-
-
-
diff --git a/docs/README_PT_BR.md b/docs/README_PT_BR.md
index 0b6aa58..4c7f617 100644
--- a/docs/README_PT_BR.md
+++ b/docs/README_PT_BR.md
@@ -139,16 +139,6 @@ MySpider().start()
TikHub.io oferece mais de 900 APIs estáveis em mais de 16 plataformas, incluindo TikTok, X, YouTube e Instagram, com mais de 40M de datasets.
Também oferece modelos de IA com desconto - Claude, GPT, GEMINI e mais com até 71% de desconto.
-
-
-
-
-
- |
-
- Nsocks fornece proxies residenciais e ISP rápidos para desenvolvedores e scrapers. Cobertura global de IPs, alto anonimato, rotação inteligente e desempenho confiável para automação e extração de dados. Use o Xcrawl para simplificar o crawling web em larga escala.
- |
-
|
@@ -170,16 +160,6 @@ MySpider().start()
Leia uma análise completa do Scrapling no The Web Scraping Club (nov. 2025), a newsletter número 1 dedicada a Web Scraping.
|
-
-
-
-
-
- |
-
- Proxies estáveis para scraping, automação e multi-accounting. IPs limpos, resposta rápida e desempenho confiável sob carga. Feito para fluxos de trabalho escaláveis.
- |
-
|
@@ -190,23 +170,38 @@ MySpider().start()
Swiftproxy fornece proxies residenciais escaláveis com mais de 80M de IPs em mais de 195 países, entregando conexões rápidas e confiáveis, rotação automática e forte desempenho anti-bloqueio. Teste grátis disponível.
|
+
+
+
+
+
+ |
+
+ 9Proxy oferece proxies residenciais a partir de apenas $0,018/IP ou $0,68/GB. Mais de 20M de IPs em mais de 90 países. Sessões fixas ou rotativas, gerenciadas pelo aplicativo desktop ou móvel.
+ |
+
+
+
+
+
+
+ |
+
+ NodeMaven - provedor de proxies confiável com a mais alta qualidade de IP do mercado. Use o código promocional SCRAPLING35 para obter 35% de desconto em proxies.
+ |
+
Quer mostrar seu anúncio aqui? Clique [aqui](https://github.com/sponsors/D4Vinci/sponsorships?tier_id=586646)
# Patrocinadores
-
-
-
-
-
diff --git a/docs/README_RU.md b/docs/README_RU.md
index dbe5b63..4dde2b9 100644
--- a/docs/README_RU.md
+++ b/docs/README_RU.md
@@ -140,16 +140,6 @@ MySpider().start()
TikHub.io предоставляет более 900 стабильных API на 16+ платформах, включая TikTok, X, YouTube и Instagram, с более чем 40 млн наборов данных.
Также предлагает AI-модели со скидкой - Claude, GPT, GEMINI и другие со скидкой до 71%.
-
-
-
-
-
- |
-
- Nsocks предоставляет быстрые резидентные и ISP прокси для разработчиков и скраперов. Глобальное покрытие IP, высокая анонимность, умная ротация и надёжная производительность для автоматизации и извлечения данных. Используйте Xcrawl для упрощения масштабного веб-краулинга.
- |
-
|
@@ -171,16 +161,6 @@ MySpider().start()
Прочитайте полный обзор Scrapling на The Web Scraping Club (ноябрь 2025) - рассылка №1, посвящённая веб-скрейпингу.
|
-
-
-
-
-
- |
-
- Стабильные прокси для скрапинга, автоматизации и мультиаккаунтинга. Чистые IP, быстрый отклик и надёжная работа под нагрузкой. Созданы для масштабируемых рабочих процессов.
- |
-
|
@@ -191,22 +171,38 @@ MySpider().start()
Swiftproxy предоставляет масштабируемые резидентные прокси с более чем 80 млн IP в 195+ странах, обеспечивая быстрые и надёжные соединения, автоматическую ротацию и высокую устойчивость к блокировкам. Доступна бесплатная пробная версия.
|
+
+
+
+
+
+ |
+
+ 9Proxy предоставляет резидентные прокси всего от $0,018 за IP или $0,68 за ГБ. Более 20 млн IP в 90+ странах. Закреплённые или ротационные сессии, управление через настольное или мобильное приложение.
+ |
+
+
+
+
+
+
+ |
+
+ NodeMaven - надёжный провайдер прокси с самым высоким качеством IP на рынке. Используйте промокод SCRAPLING35 для получения скидки 35% на прокси.
+ |
+
Хотите показать здесь свою рекламу? Нажмите [здесь](https://github.com/sponsors/D4Vinci/sponsorships?tier_id=586646)
# Спонсоры
-
-
-
-
diff --git a/docs/ai/mcp-server.md b/docs/ai/mcp-server.md
index 534c525..0ab72b5 100644
--- a/docs/ai/mcp-server.md
+++ b/docs/ai/mcp-server.md
@@ -129,7 +129,7 @@ If you are using the Docker image, then it would be something like
"ScraplingServer": {
"command": "docker",
"args": [
- "run", "-i", "--rm", "scrapling", "mcp"
+ "run", "-i", "--rm", "pyd4vinci/scrapling", "mcp"
]
}
}
diff --git a/docs/cli/extract-commands.md b/docs/cli/extract-commands.md
index 671cdcc..b3663c9 100644
--- a/docs/cli/extract-commands.md
+++ b/docs/cli/extract-commands.md
@@ -50,7 +50,7 @@ The extract command is a set of simple terminal tools that:
scrapling extract get "https://example.com" content.txt
# Or use the Docker image with something like this:
- docker run -v $(pwd)/output:/output scrapling extract get "https://blog.example.com" /output/article.md
+ docker run -v $(pwd)/output:/output pyd4vinci/scrapling extract get "https://blog.example.com" /output/article.md
```
- **Extract Specific Content**
diff --git a/docs/donate.md b/docs/donate.md
index 0228bd6..dbaab5a 100644
--- a/docs/donate.md
+++ b/docs/donate.md
@@ -1,3 +1,5 @@
+# Support and Advertisement
+
I've been creating all of these projects in my spare time and have invested considerable resources & effort in providing them to the community for free. By becoming a sponsor, you'd be directly funding my coffee reserves, helping me fulfill my responsibilities, and enabling me to continuously update existing projects and potentially create new ones.
You can sponsor me directly through the [GitHub Sponsors program](https://github.com/sponsors/D4Vinci) or [Buy Me a Coffee](https://buymeacoffee.com/d4vinci).
diff --git a/docs/index.md b/docs/index.md
index 272a86a..fdae518 100644
--- a/docs/index.md
+++ b/docs/index.md
@@ -71,25 +71,23 @@ MySpider().start()
-
-
-
-
-
-
+
+
+
+
+
+
-
-
diff --git a/images/9proxy.jpg b/images/9proxy.jpg
new file mode 100644
index 0000000..9d32ef6
Binary files /dev/null and b/images/9proxy.jpg differ
diff --git a/images/IPCook.png b/images/IPCook.png
deleted file mode 100644
index f17efc8..0000000
Binary files a/images/IPCook.png and /dev/null differ
diff --git a/images/IPFoxy.jpg b/images/IPFoxy.jpg
deleted file mode 100644
index 8d7efff..0000000
Binary files a/images/IPFoxy.jpg and /dev/null differ
diff --git a/images/MangoProxy.png b/images/MangoProxy.png
deleted file mode 100644
index 9789f04..0000000
Binary files a/images/MangoProxy.png and /dev/null differ
diff --git a/images/NodeMaven.png b/images/NodeMaven.png
new file mode 100644
index 0000000..315348c
Binary files /dev/null and b/images/NodeMaven.png differ
diff --git a/images/crawleo.png b/images/crawleo.png
deleted file mode 100644
index 7549f85..0000000
Binary files a/images/crawleo.png and /dev/null differ
diff --git a/images/nsocks.png b/images/nsocks.png
deleted file mode 100644
index 120a06c..0000000
Binary files a/images/nsocks.png and /dev/null differ
diff --git a/pyproject.toml b/pyproject.toml
index 8cb4768..9c390c0 100644
--- a/pyproject.toml
+++ b/pyproject.toml
@@ -5,7 +5,7 @@ build-backend = "setuptools.build_meta"
[project]
name = "scrapling"
# Static version instead of a dynamic version so we can get better layer caching while building docker, check the docker file to understand
-version = "0.4.8"
+version = "0.4.9"
description = "Scrapling is an undetectable, powerful, flexible, high-performance Python library that makes Web Scraping easy and effortless as it should be!"
readme = {file = "README.md", content-type = "text/markdown"}
license = {file = "LICENSE"}
@@ -61,7 +61,7 @@ classifiers = [
"Typing :: Typed",
]
dependencies = [
- "lxml>=6.1.0",
+ "lxml>=6.1.1",
"cssselect>=1.4.0",
"orjson>=3.11.8",
"tld>=0.13.2",
@@ -73,8 +73,8 @@ dependencies = [
fetchers = [
"click>=8.3.0",
"curl_cffi>=0.15.0",
- "playwright==1.59.0",
- "patchright==1.59.1",
+ "playwright==1.60.0",
+ "patchright==1.60.1",
"browserforge>=1.2.4",
"apify-fingerprint-datapoints>=0.13.0",
"msgspec>=0.21.1",
diff --git a/scrapling/__init__.py b/scrapling/__init__.py
index 3420e25..a5ff5eb 100644
--- a/scrapling/__init__.py
+++ b/scrapling/__init__.py
@@ -1,5 +1,5 @@
__author__ = "Karim Shoair (karim.shoair@pm.me)"
-__version__ = "0.4.8"
+__version__ = "0.4.9"
__copyright__ = "Copyright (c) 2024 Karim Shoair"
from typing import Any, TYPE_CHECKING
diff --git a/scrapling/cli.py b/scrapling/cli.py
index 47056d9..4b9eb86 100644
--- a/scrapling/cli.py
+++ b/scrapling/cli.py
@@ -2,6 +2,7 @@ from pathlib import Path
from subprocess import check_output
from sys import executable as python_executable
+from scrapling import __version__
from scrapling.core.utils import log
from scrapling.engines.toolbelt.custom import Response
from scrapling.core.utils._shell import _CookieParser, _ParseHeaders
@@ -10,7 +11,7 @@ from scrapling.core._types import List, Optional, Dict, Tuple, Any, Callable
from orjson import loads as json_loads, JSONDecodeError
try:
- from click import command, option, Choice, group, argument
+ from click import command, option, Choice, group, argument, version_option
except (ImportError, ModuleNotFoundError) as e:
raise ModuleNotFoundError(
"You need to install scrapling with any of the extras to enable Shell commands. See: https://scrapling.readthedocs.io/en/latest/#installation"
@@ -650,6 +651,7 @@ def stealthy_fetch(
@group()
+@version_option(version=__version__, prog_name="Scrapling")
def main():
pass
diff --git a/scrapling/engines/toolbelt/fingerprints.py b/scrapling/engines/toolbelt/fingerprints.py
index f4fdfc1..0c02002 100644
--- a/scrapling/engines/toolbelt/fingerprints.py
+++ b/scrapling/engines/toolbelt/fingerprints.py
@@ -13,8 +13,8 @@ from scrapling.core._types import Dict, Literal, Tuple
__OS_NAME__ = platform_system()
OSName = Literal["linux", "macos", "windows"]
# Current versions hardcoded for now (Playwright doesn't allow to know the version of a browser without launching it)
-chromium_version = 147
-chrome_version = 147
+chromium_version = 148
+chrome_version = 148
@lru_cache(1, typed=True)
diff --git a/scrapling/parser.py b/scrapling/parser.py
index 2b88af3..8401f27 100644
--- a/scrapling/parser.py
+++ b/scrapling/parser.py
@@ -671,7 +671,7 @@ class Selector(SelectorsGeneration):
element_data = self.retrieve(identifier or selector)
if element_data:
elements = self.relocate(element_data, percentage)
- if elements is not None and auto_save:
+ if elements and auto_save:
self.save(elements[0], identifier or selector)
return self.__handle_elements(elements)
@@ -991,7 +991,9 @@ class Selector(SelectorsGeneration):
SequenceMatcher(None, v, candidate_attributes.get(k, "")).ratio()
for k, v in original_attributes.items()
)
- checks += len(candidate_attributes)
+ # Using `max` so candidates with extra attributes are penalized and candidates
+ # with fewer attributes don't get inflated scores from a smaller denominator
+ checks += max(len(original_attributes), len(candidate_attributes))
else:
if not candidate_attributes:
# Both don't have attributes, this must mean something
diff --git a/scrapling/spiders/cache.py b/scrapling/spiders/cache.py
index 40d39d3..0305aef 100644
--- a/scrapling/spiders/cache.py
+++ b/scrapling/spiders/cache.py
@@ -64,7 +64,7 @@ class ResponseCacheManager:
async with await anyio.open_file(temp_path, "wb") as f:
await f.write(serialized)
- await temp_path.rename(self._cache_path(fingerprint))
+ await temp_path.replace(self._cache_path(fingerprint))
except Exception as e:
if await temp_path.exists():
await temp_path.unlink()
diff --git a/scrapling/spiders/checkpoint.py b/scrapling/spiders/checkpoint.py
index 25de362..95515bf 100644
--- a/scrapling/spiders/checkpoint.py
+++ b/scrapling/spiders/checkpoint.py
@@ -50,7 +50,7 @@ class CheckpointManager:
async with await anyio.open_file(temp_path, "wb") as f:
await f.write(serialized)
- await temp_path.rename(self._checkpoint_path)
+ await temp_path.replace(self._checkpoint_path)
log.info(f"Checkpoint saved: {len(data.requests)} requests, {len(data.seen)} seen URLs")
except Exception as e:
diff --git a/server.json b/server.json
index e4e813e..deccdda 100644
--- a/server.json
+++ b/server.json
@@ -14,12 +14,12 @@
"mimeType": "image/png"
}
],
- "version": "0.4.8",
+ "version": "0.4.9",
"packages": [
{
"registryType": "pypi",
"identifier": "scrapling",
- "version": "0.4.8",
+ "version": "0.4.9",
"runtimeHint": "uvx",
"packageArguments": [
{
diff --git a/setup.cfg b/setup.cfg
index 1c967c8..4e3d8b6 100644
--- a/setup.cfg
+++ b/setup.cfg
@@ -1,6 +1,6 @@
[metadata]
name = scrapling
-version = 0.4.8
+version = 0.4.9
author = Karim Shoair
author_email = karim.shoair@pm.me
description = Scrapling is an undetectable, powerful, flexible, high-performance Python library that makes Web Scraping easy and effortless as it should be!
diff --git a/tests/cli/test_cli.py b/tests/cli/test_cli.py
index 39345d0..642699d 100644
--- a/tests/cli/test_cli.py
+++ b/tests/cli/test_cli.py
@@ -4,8 +4,9 @@ from unittest.mock import patch, MagicMock
import pytest_httpbin
from scrapling.parser import Selector
+from scrapling import __version__
from scrapling.cli import (
- shell, mcp, get, post, put, delete, fetch, stealthy_fetch
+ main, shell, mcp, get, post, put, delete, fetch, stealthy_fetch
)
@@ -32,6 +33,12 @@ class TestCLI:
def runner(self):
return CliRunner()
+ def test_version_flag(self, runner):
+ """Test that the --version flag prints the Scrapling version and exits"""
+ result = runner.invoke(main, ['--version'])
+ assert result.exit_code == 0
+ assert result.output.strip() == f'Scrapling, version {__version__}'
+
def test_shell_command(self, runner):
"""Test shell command"""
with patch('scrapling.core.shell.CustomShell') as mock_shell:
diff --git a/tests/parser/test_adaptive.py b/tests/parser/test_adaptive.py
index 8d40f77..85881d1 100644
--- a/tests/parser/test_adaptive.py
+++ b/tests/parser/test_adaptive.py
@@ -56,6 +56,32 @@ class TestParserAdaptive:
assert relocated[0].has_class("new-class")
assert relocated[0].css(".new-description")[0].text == "Description 1"
+ def test_relocation_auto_save_no_match_above_threshold(self):
+ """Adaptive relocation with `auto_save=True` must not crash when no element
+ clears the `percentage` threshold (relocate() returns an empty list)."""
+ original_html = """
+
+
+ Widget
+ A widget
+
+
+ """
+ # Unrelated structure so nothing can match a high threshold
+ changed_html = "totally unrelated content"
+
+ old_page = Selector(original_html, url="example.com", adaptive=True)
+ new_page = Selector(changed_html, url="example.com", adaptive=True)
+
+ old_page.css("#target", identifier="target", auto_save=True)
+
+ # Before the fix this raised `IndexError: list index out of range` because the
+ # guard checked `elements is not None` but relocate() returns [] (never None).
+ result = new_page.css(
+ "#target", identifier="target", adaptive=True, auto_save=True, percentage=95
+ )
+ assert list(result) == []
+
@pytest.mark.asyncio
async def test_element_relocation_async(self):
"""Test relocating element after structure change in async mode"""
diff --git a/tests/parser/test_find_similar_advanced.py b/tests/parser/test_find_similar_advanced.py
index 099b0dd..95e1e69 100644
--- a/tests/parser/test_find_similar_advanced.py
+++ b/tests/parser/test_find_similar_advanced.py
@@ -2,6 +2,7 @@
Tests for Selector.find_similar() with non-default parameters.
Target file: tests/parser/test_general.py (append to TestSimilarElements class)
"""
+
import pytest
from scrapling import Selector
@@ -61,14 +62,10 @@ class TestFindSimilarAdvanced:
first = product_page.css("div.product")[0]
# Ignore both data-price and data-category → only class matters → all 3 divs match
ignore_all_data = first.find_similar(
- similarity_threshold=0.2,
- ignore_attributes=["data-price", "data-category"]
+ similarity_threshold=0.2, ignore_attributes=["data-price", "data-category"]
)
# Ignore nothing → data-category difference (fruit vs veggie) may reduce matches
- ignore_nothing = first.find_similar(
- similarity_threshold=0.9,
- ignore_attributes=[]
- )
+ ignore_nothing = first.find_similar(similarity_threshold=0.9, ignore_attributes=[])
assert len(ignore_all_data) >= len(ignore_nothing)
def test_find_similar_on_text_node_returns_empty(self, product_page):
@@ -76,3 +73,31 @@ class TestFindSimilarAdvanced:
text_node = product_page.css(".name::text")[0]
result = text_node.find_similar()
assert len(result) == 0
+
+ def test_find_similar_attribute_count_mismatch_scoring(self):
+ """The similarity denominator uses max() of both attribute counts, so candidates
+ with fewer attributes don't get inflated scores and candidates with extra
+ attributes stay penalized."""
+ html = """
+
+
+
Alpha
+
Beta
+
Gamma
+
Delta
+
+
+ """
+ page = Selector(html, adaptive=False)
+ first = page.css("div.card")[0] # Alpha
+
+ similar = first.find_similar(similarity_threshold=0.9, ignore_attributes=[])
+ texts = {el.text for el in similar}
+
+ # An exact attribute match must pass
+ assert "Delta" in texts
+ # Beta matches 1 of Alpha's 4 attributes; the old denominator counted candidate
+ # attributes only, inflating it to a perfect score (1.0 / 1)
+ assert "Beta" not in texts
+ # Gamma's extra attribute dilutes the score (4.0 / 5) - the intentional penalty
+ assert "Gamma" not in texts
diff --git a/tests/requirements.txt b/tests/requirements.txt
index 83aa3dc..7c9c3ff 100644
--- a/tests/requirements.txt
+++ b/tests/requirements.txt
@@ -1,6 +1,6 @@
pytest>=2.8.0,<9
pytest-cov
-playwright==1.59.0
+playwright==1.60.0
werkzeug<3.0.0
pytest-httpbin==2.1.0
pytest-asyncio
diff --git a/tests/spiders/test_cache.py b/tests/spiders/test_cache.py
index fc9bc59..eb1b04b 100644
--- a/tests/spiders/test_cache.py
+++ b/tests/spiders/test_cache.py
@@ -49,6 +49,27 @@ class TestResponseCacheManager:
assert dict(restored.headers) == dict(original.headers)
assert dict(restored.request_headers) == dict(original.request_headers)
+ @pytest.mark.anyio
+ async def test_put_overwrites_existing_entry(self):
+ """Re-caching the same fingerprint must replace the stored response.
+
+ Regression test for a Windows-only failure: ``Path.rename`` cannot
+ overwrite an existing destination on Windows (raising ``WinError 183``),
+ so the second ``put`` was caught by the error handler, the temp file was
+ removed, and ``get`` kept returning the stale body. ``Path.replace``
+ overwrites atomically on every platform.
+ """
+ with tempfile.TemporaryDirectory() as tmpdir:
+ cache = ResponseCacheManager(tmpdir)
+ fp = b"\x05" * 20
+
+ await cache.put(fp, _make_response(body=b"first"), "GET")
+ await cache.put(fp, _make_response(body=b"second"), "GET")
+
+ restored = await cache.get(fp)
+ assert restored is not None
+ assert restored.body == b"second"
+
@pytest.mark.anyio
async def test_get_cache_miss(self):
with tempfile.TemporaryDirectory() as tmpdir:
diff --git a/tox.ini b/tox.ini
index e740705..799ceae 100644
--- a/tox.ini
+++ b/tox.ini
@@ -10,8 +10,8 @@ envlist = pre-commit,py{310,311,312,313}
usedevelop = True
changedir = tests
deps =
- playwright==1.59.0
- patchright==1.59.1
+ playwright==1.60.0
+ patchright==1.60.1
-r{toxinidir}/tests/requirements.txt
extras = ai,shell
commands =