Guides, comparisons, and practical tutorials for reliable web data extraction.
A practical XPath cheat sheet: syntax, contains() and text matching, axes, position, and/or predicates, escaping — with copy-paste examples for scraping.
Web scraping explained: how a scraper fetches and parses a page, what companies use the data for, how it differs from crawling and APIs, and the legal lines.
The 12 business applications of web scraping, what data each one collects, and who uses it — from price monitoring to AI retrieval datasets.
Data scraping explained: how it differs from web scraping, the four main types, whether it is legal, and which tools fit each source. Written for practitioners.
Rust web scraping with current crates: parsing with scraper 0.27, bounded-concurrency crawlers, and choosing between chromiumoxide, headless_chrome, fantoccini.
Web scraping with PHP in 2026: Guzzle, Symfony DomCrawler, BrowserKit and Panther with working PHP 8 code, plus the libraries that are now abandoned.
Mechanize for web scraping in 2026: working Ruby and Python code for logins and pagination, honest maintenance status, and what to use instead of each.
Build LLM training data from the web: scrape, clean, chunk, deduplicate, and track provenance, with working Python and the legal rules that actually apply.
Current user agent strings verified July 2026, how to rotate them with matching headers and Client Hints, and where user agent rotation stops working.
Datacenter, ISP, residential, and mobile proxies for web scraping compared: how each is priced, where each breaks, and which one to reach for first.
Practical reqwest tutorial for Rust: async and blocking clients, JSON with serde, timeouts, error handling, retries, streaming, mocking, and debugging.
The Ruby web scraping gems worth using in 2026: Nokogiri, Faraday, Ferrum, Mechanize. Working code, maintenance status, and which one to pick.
Parse XML and HTML in Python: ElementTree vs lxml, XPath, namespaces, iterparse for huge files, modifying and saving documents, and XXE-safe parsing.
Compare Python web scraping libraries — requests, httpx, Beautiful Soup, lxml, Scrapy, Selenium, Playwright, curl_cffi — with a 2026 decision table.
Practical urllib3 tutorial: urllib vs urllib3 vs requests, PoolManager and connection pooling, retry strategies, timeouts, SSL verification, and proxies.
How to scrape websites with Selenium in Python: setup, selectors, waits, proxies, profiles, downloads, screenshots, exceptions, Grid, and bot detection.
How to scrape websites with PuppeteerSharp: NuGet setup, selectors, timeouts, logins, downloads, PDFs, Docker, and PuppeteerSharp vs Playwright for .NET.
How to scrape websites with Puppeteer: setup, selectors, waiting, clicks, AJAX, cookies, PDFs, proxies, stealth, Docker, and Puppeteer vs Playwright.
How to scrape websites with Playwright: install, locators, waits, headless mode, proxies, screenshots, Docker, and how it compares to Selenium and Puppeteer.
Practical Guzzle tutorial for PHP: installing, JSON and multipart requests, timeouts, error handling, retry middleware, async pools, proxies, and SSL.
Build web scrapers in n8n: HTTP Request and HTML nodes, CSS selectors, JSON parsing, pagination, scheduling, proxies, exports to Sheets/CSV, and dynamic sites.
Parse HTML in Java with jsoup: Maven setup, CSS selectors, tables, encoding, proxies, login sessions, sanitizing untrusted HTML, and Android usage.
Every JavaScript web scraping library compared: Cheerio, Puppeteer, Playwright, Crawlee. Verified 2026 versions, what each breaks on, and which are now dead.
HtmlUnit 5.x for Java web scraping: Maven setup, JDK 17 requirement, forms and sessions, JavaScript support, and when to use Playwright or an API instead.
Parse HTML in C# with Html Agility Pack: XPath and LINQ selection, class/id matching, text extraction, fixing malformed HTML, and AngleSharp comparison.
Build a TikTok scraper that survives: official APIs vs. scraping, why plain HTTP fails, and working Python code for profiles, videos, hashtags, and comments.
Real estate scraping in 2026: which portals we could actually fetch, Zillow's embedded JSON schema, working Python, and the MLS licensing rules people ignore.
We tested Indeed with datacenter, residential and stealth proxies in July 2026 — all blocked. What an Indeed scraper can still do, and what works instead.
Headless browsers explained: running Chrome and Chromium headless, essential flags, Puppeteer/Playwright/Selenium setup, memory tuning, and detection.
Copy-pasteable GPT and LLM web scraping prompts, measured token costs per page, and fixes for the failures you actually hit: bad JSON and hallucinated fields.
What Firecrawl is and how it works: getting an API key, credit pricing, self-hosting with Docker, rate limits, robots.txt, legality, and alternatives.
Every cURL command and option you need for web scraping, with copy-paste examples: GET/POST requests, redirects, cookies, auth, proxies, retries, and downloads.
Everything about HttpClient in C#/.NET with copy-paste examples: GET/POST, headers, auth, JSON, timeouts, proxies, file downloads, and IHttpClientFactory.
Verified July 2026 price-per-GB for 8 residential proxy providers, plus the minimums, expiry rules and retries that make cheap proxies expensive.
Can ChatGPT scrape websites? What ChatGPT web scraping really does, four workflows that work, and where you still need a scraping API.
Pick residential proxies for web scraping on pool size, sourcing transparency, sticky sessions, and geo granularity. 6 providers, verified July 2026 pricing.
The best proxies for web scraping, compared: verified 2026 pricing for 7 providers, a datacenter-vs-residential decision rule, and working Python code.
Web scraping with Beautiful Soup and Python: install, find() and find_all(), CSS selectors, text extraction, dynamic pages, and saving data to CSV/JSON.
AI scraping use cases that hold up in production: when LLM extraction beats CSS selectors, what it costs per page, and where it still fails.
Get started with our powerful AI-powered web scraping API. No credit card required.