---
title: "Web Scraping Blog: Guides, Tutorials & Comparisons"
description: "Web scraping guides, tutorials, and library comparisons for Python, JavaScript, Ruby, PHP, and more, plus proxy and anti-bot best practices."
url: https://webscraping.ai/blog
markdown_index: https://webscraping.ai/llms.txt
---
# Web Scraping Blog

Guides, comparisons, and practical tutorials for reliable web data extraction.

[](https://webscraping.ai/blog/web-scraping-mcp-servers)

Guides14 min read

## [Web Scraping with MCP Servers: Setup, Workflows, and Troubleshooting](https://webscraping.ai/blog/web-scraping-mcp-servers)

Set up MCP servers for web scraping in Claude, Cursor, and Claude Code: Playwright MCP, fetch and hosted servers, pagination, resources, and connection fixes.

[Read Article →](https://webscraping.ai/blog/web-scraping-mcp-servers)

[](https://webscraping.ai/blog/swift-web-scraping)

Swift12 min read

### [Swift Web Scraping: URLSession, Alamofire, SwiftSoup, and Kanna](https://webscraping.ai/blog/swift-web-scraping)

Scrape sites in Swift: URLSession vs Alamofire, JSON with Codable, HTML parsing via SwiftSoup and Kanna, redirects, WKWebView limits, and App Store rules.

[Read More →](https://webscraping.ai/blog/swift-web-scraping)

[](https://webscraping.ai/blog/scrapy-web-scraping)

Guides20 min read

### [Scrapy Web Scraping: The Complete Guide](https://webscraping.ai/blog/scrapy-web-scraping)

[Read More →](https://webscraping.ai/blog/scrapy-web-scraping)

[](https://webscraping.ai/blog/scrape-google-search-results)

Guides13 min read

### [How to Scrape Google Search Results](https://webscraping.ai/blog/scrape-google-search-results)

Why requests and BeautifulSoup no longer work on Google, a Playwright approach that does, consent screens, CAPTCHAs, SERP APIs, and the legal ground rules.

[Read More →](https://webscraping.ai/blog/scrape-google-search-results)

[](https://webscraping.ai/blog/python-requests-guide)

Python20 min read

### [Python Requests Library: The Complete Guide](https://webscraping.ai/blog/python-requests-guide)

Complete Python requests guide: GET and POST, json vs data, sessions, headers, timeouts, retries, proxies, redirects, streaming, SSL, and error handling.

[Read More →](https://webscraping.ai/blog/python-requests-guide)

[](https://webscraping.ai/blog/llm-web-scraping)

AI & ML18 min read

### [LLM Web Scraping: How AI Extraction Pipelines Actually Work](https://webscraping.ai/blog/llm-web-scraping)

[Read More →](https://webscraping.ai/blog/llm-web-scraping)

[](https://webscraping.ai/blog/http-headers-for-web-scraping)

Guides22 min read

### [HTTP Headers for Web Scraping: Anatomy of a Real Browser Request](https://webscraping.ai/blog/http-headers-for-web-scraping)

[Read More →](https://webscraping.ai/blog/http-headers-for-web-scraping)

[](https://webscraping.ai/blog/golang-web-scraping)

Go16 min read

### [Go Web Scraping: net/http, goquery, Colly, and chromedp](https://webscraping.ai/blog/golang-web-scraping)

Golang web scraping guide: net/http headers and cookies, goquery parsing, Colly crawlers, chromedp for JavaScript, and goroutine concurrency patterns.

[Read More →](https://webscraping.ai/blog/golang-web-scraping)

[](https://webscraping.ai/blog/css-selectors-cheat-sheet)

Guides20 min read

### [CSS Selectors Cheat Sheet: Syntax, Examples, and Scraping Patterns](https://webscraping.ai/blog/css-selectors-cheat-sheet)

A practical CSS selectors cheat sheet: combinators, attribute selectors, nth-child, :not and :has, text-matching limits, and per-tool syntax for web scraping.

[Read More →](https://webscraping.ai/blog/css-selectors-cheat-sheet)

[](https://webscraping.ai/blog/csharp-web-scraping)

C#15 min read

### [C# Web Scraping: The Complete Guide](https://webscraping.ai/blog/csharp-web-scraping)

Web scraping in C#: picking between HttpClient, AngleSharp, Html Agility Pack, Playwright for .NET and Selenium, with async patterns and working examples.

[Read More →](https://webscraping.ai/blog/csharp-web-scraping)

[](https://webscraping.ai/blog/cheerio-web-scraping)

JavaScript13 min read

### [Cheerio Web Scraping: Parsing HTML in Node.js](https://webscraping.ai/blog/cheerio-web-scraping)

Cheerio in Node.js: install, load HTML, CSS selectors, .each/.map loops, attributes, tables, TypeScript, plus Cheerio vs jsdom vs Puppeteer.

[Read More →](https://webscraping.ai/blog/cheerio-web-scraping)

[](https://webscraping.ai/blog/api-scraping-guide)

Guides14 min read

### [API Scraping: How to Find and Use a Website's Hidden APIs](https://webscraping.ai/blog/api-scraping-guide)

How to find a website's hidden JSON APIs with the DevTools Network tab, then handle auth, pagination, rate limits, and schema changes in production.

[Read More →](https://webscraping.ai/blog/api-scraping-guide)

[](https://webscraping.ai/blog/xpath-cheat-sheet)

Guides16 min read

### [XPath Cheat Sheet: Syntax, Functions, and Examples for Web Scraping](https://webscraping.ai/blog/xpath-cheat-sheet)

A practical XPath cheat sheet: syntax, contains() and text matching, axes, position, and/or predicates, escaping — with copy-paste examples for scraping.

[Read More →](https://webscraping.ai/blog/xpath-cheat-sheet)

[](https://webscraping.ai/blog/what-is-web-scraping-used-for)

Fundamentals4 min read

### [What is Web Scraping Used For? 12 Business Applications](https://webscraping.ai/blog/what-is-web-scraping-used-for)

The 12 business applications of web scraping, what data each one collects, and who uses it — from price monitoring to AI retrieval datasets.

[Read More →](https://webscraping.ai/blog/what-is-web-scraping-used-for)

[](https://webscraping.ai/blog/what-is-web-scraping)

Fundamentals11 min read

### [What is Web Scraping? How It Works and What It's Used For](https://webscraping.ai/blog/what-is-web-scraping)

Web scraping explained: how a scraper fetches and parses a page, what companies use the data for, how it differs from crawling and APIs, and the legal lines.

[Read More →](https://webscraping.ai/blog/what-is-web-scraping)

[](https://webscraping.ai/blog/what-is-data-scraping)

Legal10 min read

### [What is Data Scraping? Types, Tools, and Legal Limits](https://webscraping.ai/blog/what-is-data-scraping)

Data scraping explained: how it differs from web scraping, the four main types, whether it is legal, and which tools fit each source. Written for practitioners.

[Read More →](https://webscraping.ai/blog/what-is-data-scraping)

[](https://webscraping.ai/blog/web-scraping-with-rust)

Rust20 min read

### [Rust Web Scraping: the scraper Crate, TLS, PDFs, XML, and Headless Chrome](https://webscraping.ai/blog/web-scraping-with-rust)

Rust web scraping with current crates: the scraper crate, TLS certificates, XML and PDF parsing, headless Chrome, tokio concurrency, and error handling.

[Read More →](https://webscraping.ai/blog/web-scraping-with-rust)

[](https://webscraping.ai/blog/web-scraping-with-python)

Python20 min read

### [Web Scraping with Python: Complete Tutorial](https://webscraping.ai/blog/web-scraping-with-python)

[Read More →](https://webscraping.ai/blog/web-scraping-with-python)

[](https://webscraping.ai/blog/web-scraping-with-php)

PHP18 min read

### [Web Scraping with PHP: The Complete Guide](https://webscraping.ai/blog/web-scraping-with-php)

How to scrape websites with PHP: cURL and SSL, DOMDocument and XPath, Simple HTML DOM, Symfony DomCrawler, Panther for JavaScript pages, and anti-blocking.

[Read More →](https://webscraping.ai/blog/web-scraping-with-php)

[](https://webscraping.ai/blog/web-scraping-with-mechanize)

Python9 min read

### [Mechanize Web Scraping in Ruby and Python](https://webscraping.ai/blog/web-scraping-with-mechanize)

Mechanize for web scraping in 2026: working Ruby and Python code for logins and pagination, honest maintenance status, and what to use instead of each.

[Read More →](https://webscraping.ai/blog/web-scraping-with-mechanize)

[](https://webscraping.ai/blog/web-scraping-with-javascript)

JavaScript16 min read

### [Web Scraping with JavaScript and Node.js](https://webscraping.ai/blog/web-scraping-with-javascript)

[Read More →](https://webscraping.ai/blog/web-scraping-with-javascript)

[](https://webscraping.ai/blog/web-scraping-for-machine-learning)

AI & ML14 min read

### [Web Scraping for Machine Learning: Building LLM Training Data](https://webscraping.ai/blog/web-scraping-for-machine-learning)

Build LLM training data from the web: scrape, clean, chunk, deduplicate, and track provenance, with working Python and the legal rules that actually apply.

[Read More →](https://webscraping.ai/blog/web-scraping-for-machine-learning)

[](https://webscraping.ai/blog/user-agent-rotation-for-web-scraping)

Techniques12 min read

### [User Agent List for Web Scraping (2026) and How to Rotate Them](https://webscraping.ai/blog/user-agent-rotation-for-web-scraping)

Current user agent strings verified Sept 2026, how to rotate them with matching headers and Client Hints, and where user agent rotation stops working.

[Read More →](https://webscraping.ai/blog/user-agent-rotation-for-web-scraping)

[](https://webscraping.ai/blog/types-of-proxies-for-web-scraping)

Proxies10 min read

### [Proxies for Web Scraping: Types, Costs, and How to Choose](https://webscraping.ai/blog/types-of-proxies-for-web-scraping)

Datacenter, ISP, residential, and mobile proxies for web scraping compared: how each is priced, where each breaks, and which one to reach for first.

[Read More →](https://webscraping.ai/blog/types-of-proxies-for-web-scraping)

[](https://webscraping.ai/blog/rust-reqwest-guide)

JavaScript15 min read

### [Rust Reqwest Guide: HTTP Requests, JSON, Timeouts, and Retries](https://webscraping.ai/blog/rust-reqwest-guide)

Practical reqwest tutorial for Rust: async and blocking clients, JSON with serde, timeouts, error handling, retries, streaming, mocking, and debugging.

[Read More →](https://webscraping.ai/blog/rust-reqwest-guide)

[](https://webscraping.ai/blog/ruby-web-scraping-libraries)

Ruby16 min read

### [Web Scraping with Ruby: Nokogiri, HTTParty, and the Rest of the 2026 Toolkit](https://webscraping.ai/blog/ruby-web-scraping-libraries)

Web scraping with Ruby in 2026: HTTParty timeouts and SSL options, Nokogiri parsing and install fixes, Ferrum for JavaScript, PDFs, and sessions. Working code.

[Read More →](https://webscraping.ai/blog/ruby-web-scraping-libraries)

[](https://webscraping.ai/blog/python-xml-parsing)

Python16 min read

### [Python XML Parsing: ElementTree, lxml, and How to Choose](https://webscraping.ai/blog/python-xml-parsing)

Parse XML and HTML in Python: ElementTree vs lxml, XPath, namespaces, iterparse for huge files, modifying and saving documents, and XXE-safe parsing.

[Read More →](https://webscraping.ai/blog/python-xml-parsing)

[](https://webscraping.ai/blog/python-web-scraping-libraries)

Python13 min read

### [Python Web Scraping Libraries: Which One to Use in 2026](https://webscraping.ai/blog/python-web-scraping-libraries)

Compare Python web scraping libraries — requests, httpx, Beautiful Soup, lxml, Scrapy, Selenium, Playwright, curl\_cffi — with a 2026 decision table.

[Read More →](https://webscraping.ai/blog/python-web-scraping-libraries)

[](https://webscraping.ai/blog/python-urllib3-guide)

Python15 min read

### [Python urllib3 Guide: PoolManager, Retries, Timeouts, and SSL](https://webscraping.ai/blog/python-urllib3-guide)

Practical urllib3 tutorial: urllib vs urllib3 vs requests, PoolManager and connection pooling, retry strategies, timeouts, SSL verification, and proxies.

[Read More →](https://webscraping.ai/blog/python-urllib3-guide)

[](https://webscraping.ai/blog/python-selenium)

Python22 min read

### [Selenium Web Scraping: The Complete Python Guide](https://webscraping.ai/blog/python-selenium)

How to scrape websites with Selenium in Python: setup, selectors, waits, proxies, profiles, downloads, screenshots, exceptions, Grid, and bot detection.

[Read More →](https://webscraping.ai/blog/python-selenium)

[](https://webscraping.ai/blog/puppeteersharp-web-scraping)

C#13 min read

### [PuppeteerSharp Web Scraping: Headless Chrome in C#](https://webscraping.ai/blog/puppeteersharp-web-scraping)

How to scrape websites with PuppeteerSharp: NuGet setup, selectors, timeouts, logins, downloads, PDFs, Docker, and PuppeteerSharp vs Playwright for .NET.

[Read More →](https://webscraping.ai/blog/puppeteersharp-web-scraping)

[](https://webscraping.ai/blog/puppeteer-web-scraping)

JavaScript17 min read

### [Puppeteer Web Scraping: The Complete Node.js Guide](https://webscraping.ai/blog/puppeteer-web-scraping)

How to scrape websites with Puppeteer: setup, selectors, waiting, clicks, AJAX, cookies, PDFs, proxies, stealth, Docker, and Puppeteer vs Playwright.

[Read More →](https://webscraping.ai/blog/puppeteer-web-scraping)

[](https://webscraping.ai/blog/playwright-web-scraping)

Python18 min read

### [Playwright Web Scraping: The Complete Guide (Python & Node.js)](https://webscraping.ai/blog/playwright-web-scraping)

How to scrape websites with Playwright: install, locators, waits, headless mode, proxies, screenshots, Docker, and how it compares to Selenium and Puppeteer.

[Read More →](https://webscraping.ai/blog/playwright-web-scraping)

[](https://webscraping.ai/blog/php-guzzle-guide)

PHP15 min read

### [PHP Guzzle Guide: Requests, Async Concurrency, Retries, and Proxies](https://webscraping.ai/blog/php-guzzle-guide)

Practical Guzzle tutorial for PHP: installing, JSON and multipart requests, timeouts, error handling, retry middleware, async pools, proxies, and SSL.

[Read More →](https://webscraping.ai/blog/php-guzzle-guide)

[](https://webscraping.ai/blog/n8n-web-scraping)

AI & ML14 min read

### [n8n Web Scraping Guide: HTTP Requests, HTML Extraction, and Browser Automation](https://webscraping.ai/blog/n8n-web-scraping)

Build web scrapers in n8n: HTTP Request and HTML nodes, CSS selectors, JSON parsing, pagination, scheduling, proxies, exports to Sheets/CSV, and dynamic sites.

[Read More →](https://webscraping.ai/blog/n8n-web-scraping)

[](https://webscraping.ai/blog/jsoup-java-guide)

JavaScript18 min read

### [jsoup Java Web Scraping: HTML Parsing and CSS Selectors](https://webscraping.ai/blog/jsoup-java-guide)

Java web scraping with jsoup: Maven setup, CSS selectors, tables, encoding, proxies and country targeting, caching, login sessions, and JavaScript pages.

[Read More →](https://webscraping.ai/blog/jsoup-java-guide)

[](https://webscraping.ai/blog/javascript-web-scraping-libraries)

JavaScript12 min read

### [JavaScript Web Scraping Libraries: What to Use in 2026](https://webscraping.ai/blog/javascript-web-scraping-libraries)

Every JavaScript web scraping library compared: Cheerio, Puppeteer, Playwright, Crawlee. Verified 2026 versions, what each breaks on, and which are now dead.

[Read More →](https://webscraping.ai/blog/javascript-web-scraping-libraries)

[](https://webscraping.ai/blog/is-web-scraping-legal)

Legal14 min read

### [Is Web Scraping Legal? What the Case Law Actually Says](https://webscraping.ai/blog/is-web-scraping-legal)

[Read More →](https://webscraping.ai/blog/is-web-scraping-legal)

[](https://webscraping.ai/blog/htmlunit-web-scraping-java)

Java11 min read

### [HtmlUnit Web Scraping in Java: Setup, Forms, and Limits](https://webscraping.ai/blog/htmlunit-web-scraping-java)

HtmlUnit 5.x for Java web scraping: Maven setup, JDK 17 requirement, forms and sessions, JavaScript support, and when to use Playwright or an API instead.

[Read More →](https://webscraping.ai/blog/htmlunit-web-scraping-java)

[](https://webscraping.ai/blog/html-agility-pack-csharp)

C#14 min read

### [Html Agility Pack: C# HTML Parsing and Web Scraping Guide](https://webscraping.ai/blog/html-agility-pack-csharp)

Parse HTML in C# with Html Agility Pack: XPath and LINQ selection, class/id matching, text extraction, fixing malformed HTML, and AngleSharp comparison.

[Read More →](https://webscraping.ai/blog/html-agility-pack-csharp)

[](https://webscraping.ai/blog/how-to-scrape-tiktok)

Guides12 min read

### [TikTok Scraper: How to Scrape TikTok Data in 2026](https://webscraping.ai/blog/how-to-scrape-tiktok)

Build a TikTok scraper that survives: official APIs vs. scraping, why plain HTTP fails, and working Python code for profiles, videos, hashtags, and comments.

[Read More →](https://webscraping.ai/blog/how-to-scrape-tiktok)

[](https://webscraping.ai/blog/how-to-scrape-real-estate-data)

Guides10 min read

### [Real Estate Scraping: How to Collect Property Data in 2026](https://webscraping.ai/blog/how-to-scrape-real-estate-data)

Real estate scraping in 2026: which portals we could actually fetch, Zillow's embedded JSON schema, working Python, and the MLS licensing rules people ignore.

[Read More →](https://webscraping.ai/blog/how-to-scrape-real-estate-data)

[](https://webscraping.ai/blog/how-to-scrape-indeed)

Guides9 min read

### [How to Scrape Indeed in 2026: What an Indeed Scraper Can and Can't Do](https://webscraping.ai/blog/how-to-scrape-indeed)

We tested Indeed with datacenter, residential and stealth proxies in July 2026 — all blocked. What an Indeed scraper can still do, and what works instead.

[Read More →](https://webscraping.ai/blog/how-to-scrape-indeed)

[](https://webscraping.ai/blog/headless-browser-guide)

Fundamentals17 min read

### [What Is a Headless Browser? Headless Chrome for Scraping and Testing](https://webscraping.ai/blog/headless-browser-guide)

Headless browsers explained: running Chrome and Chromium headless, essential flags, Puppeteer/Playwright/Selenium setup, memory tuning, and detection.

[Read More →](https://webscraping.ai/blog/headless-browser-guide)

[](https://webscraping.ai/blog/gpt-prompts-for-web-scraping)

Guides12 min read

### [GPT and LLM Web Scraping Prompts That Hold Up in Production](https://webscraping.ai/blog/gpt-prompts-for-web-scraping)

Copy-pasteable GPT and LLM web scraping prompts, measured token costs per page, and fixes for the failures you actually hit: bad JSON and hallucinated fields.

[Read More →](https://webscraping.ai/blog/gpt-prompts-for-web-scraping)

[](https://webscraping.ai/blog/firecrawl-guide)

Guides13 min read

### [Firecrawl Guide: Self-Hosting, Pricing, API Keys, and Limits](https://webscraping.ai/blog/firecrawl-guide)

What Firecrawl is and how it works: getting an API key, credit pricing, self-hosting with Docker, rate limits, robots.txt, legality, and alternatives.

[Read More →](https://webscraping.ai/blog/firecrawl-guide)

[](https://webscraping.ai/blog/curl-commands-for-web-scraping)

Guides16 min read

### [cURL Commands and Options for Web Scraping: Complete Guide](https://webscraping.ai/blog/curl-commands-for-web-scraping)

Every cURL command and option you need for web scraping, with copy-paste examples: GET/POST requests, redirects, cookies, auth, proxies, retries, and downloads.

[Read More →](https://webscraping.ai/blog/curl-commands-for-web-scraping)

[](https://webscraping.ai/blog/csharp-httpclient-guide)

C#17 min read

### [C# HttpClient: The Complete Guide with Examples](https://webscraping.ai/blog/csharp-httpclient-guide)

Everything about HttpClient in C#/.NET with copy-paste examples: GET/POST, headers, auth, JSON, timeouts, proxies, file downloads, and IHttpClientFactory.

[Read More →](https://webscraping.ai/blog/csharp-httpclient-guide)

[](https://webscraping.ai/blog/cheapest-residential-proxies)

Proxies10 min read

### [Cheapest Residential Proxies in 2026: Verified Price per GB](https://webscraping.ai/blog/cheapest-residential-proxies)

Verified Sept 2026 price-per-GB for 9 residential proxy providers, plus the minimums, expiry rules and retries that make cheap proxies expensive.

[Read More →](https://webscraping.ai/blog/cheapest-residential-proxies)

[](https://webscraping.ai/blog/chatgpt-use-cases)

Guides7 min read

### [ChatGPT Web Scraping: Use Cases That Actually Work](https://webscraping.ai/blog/chatgpt-use-cases)

Can ChatGPT scrape websites? What ChatGPT web scraping really does, four workflows that work, and where you still need a scraping API.

[Read More →](https://webscraping.ai/blog/chatgpt-use-cases)

[](https://webscraping.ai/blog/best-residential-proxy-providers-for-web-scraping)

Proxies8 min read

### [Best Residential Proxies for Web Scraping: 6 Providers Compared](https://webscraping.ai/blog/best-residential-proxy-providers-for-web-scraping)

Pick residential proxies for web scraping on pool size, sourcing transparency, sticky sessions, and geo granularity. 6 providers, September 2026 pricing.

[Read More →](https://webscraping.ai/blog/best-residential-proxy-providers-for-web-scraping)

[](https://webscraping.ai/blog/best-proxy-providers-for-web-scraping)

Proxies10 min read

### [Best Proxies for Web Scraping: 7 Providers Compared (2026)](https://webscraping.ai/blog/best-proxy-providers-for-web-scraping)

The best proxies for web scraping, compared: verified 2026 pricing for 7 providers, a datacenter-vs-residential decision rule, and working Python code.

[Read More →](https://webscraping.ai/blog/best-proxy-providers-for-web-scraping)

[](https://webscraping.ai/blog/best-free-proxy-lists)

Proxies13 min read

### [Free Proxy List APIs and Scrapers for Web Scraping](https://webscraping.ai/blog/best-free-proxy-lists)

[Read More →](https://webscraping.ai/blog/best-free-proxy-lists)

[](https://webscraping.ai/blog/best-datacenter-proxy-providers-for-web-scraping)

Proxies11 min read

### [Best Datacenter Proxies for Web Scraping: 6 Providers Compared](https://webscraping.ai/blog/best-datacenter-proxy-providers-for-web-scraping)

[Read More →](https://webscraping.ai/blog/best-datacenter-proxy-providers-for-web-scraping)

[](https://webscraping.ai/blog/beautiful-soup-python)

Python18 min read

### [Beautiful Soup Web Scraping in Python: The Complete Guide](https://webscraping.ai/blog/beautiful-soup-python)

Web scraping with Beautiful Soup and Python: install, find() and find\_all(), CSS selectors, text extraction, dynamic pages, and saving data to CSV/JSON.

[Read More →](https://webscraping.ai/blog/beautiful-soup-python)

[](https://webscraping.ai/blog/ai-use-cases)

AI & ML8 min read

### [AI Scraping Use Cases: Where AI Data Extraction Pays Off](https://webscraping.ai/blog/ai-use-cases)

AI scraping use cases that hold up in production: when LLM extraction beats CSS selectors, what it costs per page, and where it still fails.

[Read More →](https://webscraping.ai/blog/ai-use-cases)

## Ready to Start Web Scraping?

Get started with our powerful AI-powered web scraping API. No credit card required.

[Start Free Trial](https://webscraping.ai/auth/sign_up)
