Everything you need to integrate WebScraping.AI into your applications.
Updated
Get started with WebScraping.AI in under 5 minutes. Here's a simple example to extract data from any webpage.
Sign up at webscraping.ai to get your free API key with 2,000 credits.
Try the AI question endpoint to ask a question about any webpage:
curl -G "https://api.webscraping.ai/ai/question" \
--data-urlencode "api_key=YOUR_API_KEY" \
--data-urlencode "url=https://example.com" \
--data-urlencode "question=What is this page about?"
# pip install webscraping_ai
# https://pypi.org/project/webscraping-ai/
from webscraping_ai import Client
client = Client(api_key="YOUR_API_KEY")
answer = client.question("https://example.com", question="What is this page about?")
print(answer)
// npm install webscraping-ai
// https://www.npmjs.com/package/webscraping-ai
import { WebScrapingAI } from 'webscraping-ai';
const client = new WebScrapingAI({ apiKey: 'YOUR_API_KEY' });
const answer = await client.question({
url: 'https://example.com',
question: 'What is this page about?',
});
console.log(answer);
<?php
// composer require webscraping-ai/webscraping-ai-php
// https://packagist.org/packages/webscraping-ai/webscraping-ai-php
require 'vendor/autoload.php';
use WebScrapingAI\Client;
$client = new Client('YOUR_API_KEY');
$answer = $client->question('https://example.com', 'What is this page about?');
echo $answer;
# gem install webscraping_ai
# https://rubygems.org/gems/webscraping_ai
require 'webscraping_ai'
client = WebScrapingAI::Client.new(api_key: 'YOUR_API_KEY')
answer = client.question('https://example.com', question: 'What is this page about?')
puts answer
// go get github.com/webscraping-ai/webscraping-ai-go/v4
// https://pkg.go.dev/github.com/webscraping-ai/webscraping-ai-go/v4
package main
import (
"context"
"fmt"
webscrapingai "github.com/webscraping-ai/webscraping-ai-go/v4"
)
func main() {
client, _ := webscrapingai.NewClient(&webscrapingai.Config{APIKey: "YOUR_API_KEY"})
answer, _ := client.Question(context.Background(), &webscrapingai.QuestionOptions{
URL: "https://example.com",
Question: "What is this page about?",
})
fmt.Println(answer)
}
// Maven: ai.webscraping:webscraping-ai:4.2.0
// https://central.sonatype.com/artifact/ai.webscraping/webscraping-ai
import ai.webscraping.Client;
import ai.webscraping.Config;
import ai.webscraping.option.QuestionOptions;
Client client = new Client(Config.builder().apiKey("YOUR_API_KEY").build());
String answer = client.question(QuestionOptions.builder()
.url("https://example.com")
.question("What is this page about?")
.build());
System.out.println(answer);
// dotnet add package WebScrapingAI
// https://www.nuget.org/packages/WebScrapingAI
using WebScrapingAI;
var client = new WebScrapingAIClient(new WebScrapingAIClientOptions { ApiKey = "YOUR_API_KEY" });
var answer = await client.QuestionAsync(new QuestionRequest {
Url = "https://example.com",
Question = "What is this page about?",
});
Console.WriteLine(answer);
The API returns the AI-generated answer as plain text:
All API requests require an API key. Pass your key as a query parameter:
GET https://api.webscraping.ai/html?url=https://example.com&api_key=YOUR_API_KEY
All API endpoints use this base URL:
/serp) are a flat 15 credits per request and structured data (/data) is 15 credits per request (50 for Reddit) — see the table below.timeout parameter to control duration./data, pages that come back parse_failed or not_found count as successful and are charged; only failed fetches are free.| Configuration | Credits | Notes |
|---|---|---|
| Basic (no JS, datacenter proxy) | 1 | Fastest, for static sites |
| With JS rendering | 5 | Default setting, headless Chrome |
| Residential proxy (no JS) | 10 | For anti-bot protected sites |
| Residential proxy + JS | 25 | Maximum compatibility |
| Stealth proxy | 50 | For sites with the strongest anti-bot protection |
| AI endpoints (/ai/question, /ai/fields) | +5 | Added on top of the proxy/JS cost above |
| Search results (/serp) | 15 | Flat per search; proxy/JS settings don't apply |
| Structured data (/data) | 15 (Reddit: 50) | Per page, including pages that parse empty or don't exist; proxy/JS settings don't apply |
All costs are per successful request. Failed requests are free.
Most failed requests succeed on a retry or with a different configuration. If you encounter failures:
timeout to 20000-25000ms for slow-loading websitesproxy=residential if datacenter proxies are blockedjs_timeout for pages with slow-loading dynamic contentJavaScript Rendering
JS rendering is enabled by default (js=true) using headless Chrome. Keep enabled for SPAs (React, Vue, Angular), AJAX content, and dynamic pages. Disable (js=false) for static sites, faster responses, or server-rendered content.
Proxy Strategy
Start with datacenter proxies (default) for speed and cost. Switch to residential proxies if: website blocks datacenter IPs, getting 403 errors, need to bypass anti-bot protection, or scraping geo-restricted content.
Use our official SDKs for easier integration in your preferred language.
pip install webscraping_ai
npm install webscraping-ai
composer require webscraping-ai/webscraping-ai-php
gem install webscraping_ai
go get github.com/webscraping-ai/webscraping-ai-go/v4
ai.webscraping:webscraping-ai
dotnet add package WebScrapingAI
https://mcp.webscraping.ai/mcp
npm install -g webscraping-ai-cli
n8n-nodes-webscraping-ai
Need an SDK for another language? Let us know!
Integrate WebScraping.AI with AI assistants and LLM platforms.
Our hosted MCP server integrates WebScraping.AI directly with AI assistants that support the Model Context Protocol — Claude, Claude Desktop, Claude Code, Cursor, Codex, Windsurf, and any MCP-compatible platform. Add the URL to your client and sign in with your WebScraping.AI account; no API key needed:
# Remote MCP server URL (Streamable HTTP, OAuth login)
https://mcp.webscraping.ai/mcp
# Example: Claude Code
claude mcp add --transport http webscraping-ai https://mcp.webscraping.ai/mcpPrefer a self-hosted or customizable server? The previous open-source npm version runs locally over stdio with your API key (npx -y webscraping-ai-mcp) and exposes the same 9 tools, including webscraping_ai_serp for Google search results and webscraping_ai_data for structured data from supported sites. See the setup guide for both options.
Our command-line tool wraps every endpoint for use in scripts and terminals, and ships an AI agent skill your coding assistant can install:
# Install the CLI
npm install -g webscraping-ai-cli
# Google search results for a query (alias: search)
webscraping-ai serp coffee machines --gl us --hl en --page 1
# Structured data for a page on a supported site
webscraping-ai data 'https://www.youtube.com/watch?v=dQw4w9WgXcQ'
# Install the agent skill into your editor(s)
webscraping-ai setup skill --allThe skill teaches AI coding assistants — Claude Code, Cursor, Windsurf, Kiro, OpenCode, Gemini CLI, GitHub Copilot, Augment, and Factory — when and how to fetch live page content through the CLI.
Use our OpenAPI specification (also available as JSON) to integrate with AI tools that support API schemas, such as GPT Actions or custom agents.
Point your coding agent or AI assistant at these instead of the HTML pages:
.md to its path (/docs.md) or request it with Accept: text/markdown. Each docs section has its own file, e.g. /docs/ai-fields.md.Use WebScraping.AI as a proxy server for your existing tools. Route requests through our infrastructure without changing your code.
| Host | proxy.webscraping.ai |
| Port | 8888 |
| Username | Your API key |
| Password | Parameters (e.g., js=true&proxy=residential) |
Proxy Mode is designed for tools you already use — HTTP clients, headless browsers, scraping frameworks. The snippets below show the most common integrations. For a more idiomatic, code-first integration, see our official SDKs instead.
# https://curl.se/docs/manpage.html#-x
curl -x "http://YOUR_API_KEY:js=true&proxy=residential@proxy.webscraping.ai:8888" \
-k "https://example.com"# pip install requests
# https://pypi.org/project/requests/
import requests
proxy_url = "http://YOUR_API_KEY:js=true&proxy=residential@proxy.webscraping.ai:8888"
response = requests.get(
"https://example.com",
proxies={"http": proxy_url, "https": proxy_url},
verify=False, # proxy uses a self-signed cert
)
print(response.text)# pip install scrapy
# https://docs.scrapy.org/en/latest/topics/downloader-middleware.html#module-scrapy.downloadermiddlewares.httpproxy
import scrapy
PROXY = "http://YOUR_API_KEY:js=true&proxy=residential@proxy.webscraping.ai:8888"
class ExampleSpider(scrapy.Spider):
name = "example"
# Scrapy doesn't verify TLS certificates by default,
# so the proxy's self-signed cert needs no extra setting
def start_requests(self):
yield scrapy.Request(
"https://example.com",
meta={"proxy": PROXY},
callback=self.parse,
)
def parse(self, response):
yield {"title": response.css("title::text").get()}# pip install selenium-wire "blinker<1.8"
# https://pypi.org/project/selenium-wire/ (vanilla Selenium can't pass proxy credentials;
# selenium-wire is archived and breaks with blinker 1.8+, hence the pin)
from seleniumwire import webdriver
PROXY = "http://YOUR_API_KEY:js=true&proxy=residential@proxy.webscraping.ai:8888"
seleniumwire_options = {
"proxy": {"http": PROXY, "https": PROXY, "no_proxy": "localhost,127.0.0.1"},
"verify_ssl": False,
}
driver = webdriver.Chrome(seleniumwire_options=seleniumwire_options)
driver.get("https://example.com")
print(driver.page_source)
driver.quit()// npm install playwright
// https://playwright.dev/docs/network#http-proxy
const { chromium } = require('playwright');
(async () => {
const browser = await chromium.launch({
proxy: {
server: 'http://proxy.webscraping.ai:8888',
username: 'YOUR_API_KEY',
password: 'js=true&proxy=residential',
},
});
const context = await browser.newContext({ ignoreHTTPSErrors: true });
const page = await context.newPage();
await page.goto('https://example.com');
console.log(await page.content());
await browser.close();
})();// npm install puppeteer
// https://pptr.dev/api/puppeteer.page.authenticate
const puppeteer = require('puppeteer');
(async () => {
const browser = await puppeteer.launch({
args: [
'--proxy-server=proxy.webscraping.ai:8888',
'--ignore-certificate-errors',
],
});
const page = await browser.newPage();
await page.authenticate({
username: 'YOUR_API_KEY',
password: 'js=true&proxy=residential',
});
await page.goto('https://example.com');
console.log(await page.content());
await browser.close();
})();-k in cURL, verify=False in requests, verify_ssl: False in selenium-wire, ignoreHTTPSErrors: true in Playwright, --ignore-certificate-errors in Puppeteer (Scrapy doesn't verify certificates by default).
Use AI to answer questions about any webpage. Perfect for extracting specific information without parsing HTML.
/ai/question
| Parameter | Type | Description |
|---|---|---|
| url required | string | URL of the webpage to analyze |
| question required | string | Question to ask about the page content |
| api_key required | string | Your API key |
curl -G "https://api.webscraping.ai/ai/question" \
--data-urlencode "api_key=YOUR_API_KEY" \
--data-urlencode "url=https://news.ycombinator.com" \
--data-urlencode "question=What are the top 3 stories on this page?"# pip install webscraping_ai
# https://pypi.org/project/webscraping-ai/
from webscraping_ai import Client
client = Client(api_key="YOUR_API_KEY")
answer = client.question(
"https://news.ycombinator.com",
question="What are the top 3 stories on this page?",
)
print(answer)// npm install webscraping-ai
// https://www.npmjs.com/package/webscraping-ai
import { WebScrapingAI } from 'webscraping-ai';
const client = new WebScrapingAI({ apiKey: 'YOUR_API_KEY' });
const answer = await client.question({
url: 'https://news.ycombinator.com',
question: 'What are the top 3 stories on this page?',
});
console.log(answer);<?php
// composer require webscraping-ai/webscraping-ai-php
// https://packagist.org/packages/webscraping-ai/webscraping-ai-php
require 'vendor/autoload.php';
use WebScrapingAI\Client;
$client = new Client('YOUR_API_KEY');
$answer = $client->question(
'https://news.ycombinator.com',
'What are the top 3 stories on this page?',
);
echo $answer;# gem install webscraping_ai
# https://rubygems.org/gems/webscraping_ai
require 'webscraping_ai'
client = WebScrapingAI::Client.new(api_key: 'YOUR_API_KEY')
answer = client.question(
'https://news.ycombinator.com',
question: 'What are the top 3 stories on this page?'
)
puts answer// go get github.com/webscraping-ai/webscraping-ai-go/v4
// https://pkg.go.dev/github.com/webscraping-ai/webscraping-ai-go/v4
package main
import (
"context"
"fmt"
webscrapingai "github.com/webscraping-ai/webscraping-ai-go/v4"
)
func main() {
client, _ := webscrapingai.NewClient(&webscrapingai.Config{APIKey: "YOUR_API_KEY"})
answer, _ := client.Question(context.Background(), &webscrapingai.QuestionOptions{
URL: "https://news.ycombinator.com",
Question: "What are the top 3 stories on this page?",
})
fmt.Println(answer)
}// Maven: ai.webscraping:webscraping-ai:4.2.0
// https://central.sonatype.com/artifact/ai.webscraping/webscraping-ai
import ai.webscraping.Client;
import ai.webscraping.Config;
import ai.webscraping.option.QuestionOptions;
Client client = new Client(Config.builder().apiKey("YOUR_API_KEY").build());
String answer = client.question(QuestionOptions.builder()
.url("https://news.ycombinator.com")
.question("What are the top 3 stories on this page?")
.build());
System.out.println(answer);// dotnet add package WebScrapingAI
// https://www.nuget.org/packages/WebScrapingAI
using WebScrapingAI;
var client = new WebScrapingAIClient(new WebScrapingAIClientOptions { ApiKey = "YOUR_API_KEY" });
var answer = await client.QuestionAsync(new QuestionRequest {
Url = "https://news.ycombinator.com",
Question = "What are the top 3 stories on this page?",
});
Console.WriteLine(answer);Extract specific data fields from any webpage as structured JSON. Ideal for scraping product details, articles, profiles, and more.
/ai/fields
| Parameter | Type | Description |
|---|---|---|
| url required | string | URL of the webpage to extract from |
| fields required | object | Object with field names as keys and extraction instructions as values |
| api_key required | string | Your API key |
curl -G "https://api.webscraping.ai/ai/fields" \
--data-urlencode "api_key=YOUR_API_KEY" \
--data-urlencode "url=https://amazon.com/dp/B08N5WRWNW" \
--data-urlencode "fields[title]=Product title" \
--data-urlencode "fields[price]=Current price with currency" \
--data-urlencode "fields[rating]=Average star rating"# pip install webscraping_ai
# https://pypi.org/project/webscraping-ai/
from webscraping_ai import Client
client = Client(api_key="YOUR_API_KEY")
result = client.fields(
"https://amazon.com/dp/B08N5WRWNW",
fields={
"title": "Product title",
"price": "Current price with currency",
"rating": "Average star rating",
},
)
print(result)// npm install webscraping-ai
// https://www.npmjs.com/package/webscraping-ai
import { WebScrapingAI } from 'webscraping-ai';
const client = new WebScrapingAI({ apiKey: 'YOUR_API_KEY' });
const result = await client.fields({
url: 'https://amazon.com/dp/B08N5WRWNW',
fields: {
title: 'Product title',
price: 'Current price with currency',
rating: 'Average star rating',
},
});
console.log(result);<?php
// composer require webscraping-ai/webscraping-ai-php
// https://packagist.org/packages/webscraping-ai/webscraping-ai-php
require 'vendor/autoload.php';
use WebScrapingAI\Client;
$client = new Client('YOUR_API_KEY');
$result = $client->fields('https://amazon.com/dp/B08N5WRWNW', [
'title' => 'Product title',
'price' => 'Current price with currency',
'rating' => 'Average star rating',
]);
print_r($result);# gem install webscraping_ai
# https://rubygems.org/gems/webscraping_ai
require 'webscraping_ai'
client = WebScrapingAI::Client.new(api_key: 'YOUR_API_KEY')
result = client.fields(
'https://amazon.com/dp/B08N5WRWNW',
fields: {
title: 'Product title',
price: 'Current price with currency',
rating: 'Average star rating'
}
)
puts result.inspect// go get github.com/webscraping-ai/webscraping-ai-go/v4
// https://pkg.go.dev/github.com/webscraping-ai/webscraping-ai-go/v4
package main
import (
"context"
"fmt"
webscrapingai "github.com/webscraping-ai/webscraping-ai-go/v4"
)
func main() {
client, _ := webscrapingai.NewClient(&webscrapingai.Config{APIKey: "YOUR_API_KEY"})
result, _ := client.Fields(context.Background(), &webscrapingai.FieldsOptions{
URL: "https://amazon.com/dp/B08N5WRWNW",
Fields: map[string]string{
"title": "Product title",
"price": "Current price with currency",
"rating": "Average star rating",
},
})
fmt.Println(result.Result)
}// Maven: ai.webscraping:webscraping-ai:4.2.0
// https://central.sonatype.com/artifact/ai.webscraping/webscraping-ai
import ai.webscraping.Client;
import ai.webscraping.Config;
import ai.webscraping.option.FieldsOptions;
import ai.webscraping.result.FieldsResult;
Client client = new Client(Config.builder().apiKey("YOUR_API_KEY").build());
FieldsResult result = client.fields(FieldsOptions.builder()
.url("https://amazon.com/dp/B08N5WRWNW")
.addField("title", "Product title")
.addField("price", "Current price with currency")
.addField("rating", "Average star rating")
.build());
System.out.println(result.getResult());// dotnet add package WebScrapingAI
// https://www.nuget.org/packages/WebScrapingAI
using WebScrapingAI;
var client = new WebScrapingAIClient(new WebScrapingAIClientOptions { ApiKey = "YOUR_API_KEY" });
var result = await client.FieldsAsync(new FieldsRequest {
Url = "https://amazon.com/dp/B08N5WRWNW",
Fields = new Dictionary<string, string> {
["title"] = "Product title",
["price"] = "Current price with currency",
["rating"] = "Average star rating",
},
});
Console.WriteLine(result.Result);The extracted fields come back under a result key; a field the page doesn't contain is null.
Fetch the full HTML content of any webpage. Includes JavaScript rendering via headless Chrome and automatic proxy rotation.
/html
| Parameter | Type | Description |
|---|---|---|
| url required | string | URL of the webpage to fetch |
| api_key required | string | Your API key |
| js optional | boolean | Enable JavaScript rendering (default: true) |
| proxy optional | string | Proxy type: "datacenter" (default), "residential", "stealth" or "auto" |
curl -G "https://api.webscraping.ai/html" \
--data-urlencode "api_key=YOUR_API_KEY" \
--data-urlencode "url=https://example.com" \
--data-urlencode "js=true"# pip install webscraping_ai
# https://pypi.org/project/webscraping-ai/
from webscraping_ai import Client
client = Client(api_key="YOUR_API_KEY")
html = client.html("https://example.com", js=True)
print(html)// npm install webscraping-ai
// https://www.npmjs.com/package/webscraping-ai
import { WebScrapingAI } from 'webscraping-ai';
const client = new WebScrapingAI({ apiKey: 'YOUR_API_KEY' });
const html = await client.html({ url: 'https://example.com', js: true });
console.log(html);<?php
// composer require webscraping-ai/webscraping-ai-php
// https://packagist.org/packages/webscraping-ai/webscraping-ai-php
require 'vendor/autoload.php';
use WebScrapingAI\Client;
$client = new Client('YOUR_API_KEY');
$html = $client->html('https://example.com', js: true);
echo $html;# gem install webscraping_ai
# https://rubygems.org/gems/webscraping_ai
require 'webscraping_ai'
client = WebScrapingAI::Client.new(api_key: 'YOUR_API_KEY')
html = client.html('https://example.com', js: true)
puts html// go get github.com/webscraping-ai/webscraping-ai-go/v4
// https://pkg.go.dev/github.com/webscraping-ai/webscraping-ai-go/v4
package main
import (
"context"
"fmt"
webscrapingai "github.com/webscraping-ai/webscraping-ai-go/v4"
)
func main() {
client, _ := webscrapingai.NewClient(&webscrapingai.Config{APIKey: "YOUR_API_KEY"})
js := true
html, _ := client.HTML(context.Background(), &webscrapingai.HTMLOptions{
URL: "https://example.com",
CommonOptions: webscrapingai.CommonOptions{JS: &js},
})
fmt.Println(html)
}// Maven: ai.webscraping:webscraping-ai:4.2.0
// https://central.sonatype.com/artifact/ai.webscraping/webscraping-ai
import ai.webscraping.Client;
import ai.webscraping.Config;
import ai.webscraping.option.HtmlOptions;
Client client = new Client(Config.builder().apiKey("YOUR_API_KEY").build());
String html = client.html(HtmlOptions.builder()
.url("https://example.com")
.js(true)
.build());
System.out.println(html);// dotnet add package WebScrapingAI
// https://www.nuget.org/packages/WebScrapingAI
using WebScrapingAI;
var client = new WebScrapingAIClient(new WebScrapingAIClientOptions { ApiKey = "YOUR_API_KEY" });
var html = await client.HtmlAsync(new HtmlRequest {
Url = "https://example.com",
Js = true,
});
Console.WriteLine(html);/html
Use POST requests to send data to the target page (e.g., form submissions, API calls). The /text, /selected and /selected-multiple endpoints accept POST the same way.
| query string | params | api_key, url and all other parameters go in the query string, as with GET |
| request body | raw | The body of your POST request is sent to the target URL, with your request's Content-Type. JSON and text/plain bodies pass through as-is; to send a form-encoded body, post it as text/plain and set headers[Content-Type]=application/x-www-form-urlencoded |
# Send a JSON body to a target page
curl -X POST "https://api.webscraping.ai/html?api_key=YOUR_API_KEY&url=https%3A%2F%2Fhttpbin.org%2Fpost" \
-H "Content-Type: application/json" \
-d '{"username": "test", "password": "demo"}'import requests
# Send a JSON body to a target page
response = requests.post(
"https://api.webscraping.ai/html",
params={"api_key": "YOUR_API_KEY", "url": "https://httpbin.org/post"},
json={"username": "test", "password": "demo"},
)
print(response.text)// Send a JSON body to a target page
const query = new URLSearchParams({
api_key: 'YOUR_API_KEY',
url: 'https://httpbin.org/post',
});
const response = await fetch(`https://api.webscraping.ai/html?${query}`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ username: 'test', password: 'demo' }),
});
console.log(await response.text());Convert a webpage to clean Markdown — boilerplate stripped, headings, lists, and links preserved (table cells come through as plain text). Perfect for feeding content to LLMs and RAG pipelines. Try it on any URL.
/text
| Parameter | Type | Description |
|---|---|---|
| url required | string | URL of the webpage |
| text_format optional | string | "plain" (default) returns raw Markdown; "json"/"xml" wrap it with title and description |
| return_links optional | boolean | Include links in JSON response (default: false) |
curl -G "https://api.webscraping.ai/text" \
--data-urlencode "api_key=YOUR_API_KEY" \
--data-urlencode "url=https://example.com" \
--data-urlencode "text_format=json"# pip install webscraping_ai
# https://pypi.org/project/webscraping-ai/
from webscraping_ai import Client
client = Client(api_key="YOUR_API_KEY")
result = client.text("https://example.com", text_format="json")
print(result)// npm install webscraping-ai
// https://www.npmjs.com/package/webscraping-ai
import { WebScrapingAI } from 'webscraping-ai';
const client = new WebScrapingAI({ apiKey: 'YOUR_API_KEY' });
const result = await client.text({
url: 'https://example.com',
text_format: 'json',
});
console.log(result);<?php
// composer require webscraping-ai/webscraping-ai-php
// https://packagist.org/packages/webscraping-ai/webscraping-ai-php
require 'vendor/autoload.php';
use WebScrapingAI\Client;
$client = new Client('YOUR_API_KEY');
$result = $client->text('https://example.com', textFormat: 'json');
print_r($result);# gem install webscraping_ai
# https://rubygems.org/gems/webscraping_ai
require 'webscraping_ai'
client = WebScrapingAI::Client.new(api_key: 'YOUR_API_KEY')
result = client.text('https://example.com', text_format: 'json')
puts result.inspect// go get github.com/webscraping-ai/webscraping-ai-go/v4
// https://pkg.go.dev/github.com/webscraping-ai/webscraping-ai-go/v4
package main
import (
"context"
"fmt"
webscrapingai "github.com/webscraping-ai/webscraping-ai-go/v4"
)
func main() {
client, _ := webscrapingai.NewClient(&webscrapingai.Config{APIKey: "YOUR_API_KEY"})
text, _ := client.Text(context.Background(), &webscrapingai.TextOptions{
URL: "https://example.com",
TextFormat: "json",
})
fmt.Println(text)
}// Maven: ai.webscraping:webscraping-ai:4.2.0
// https://central.sonatype.com/artifact/ai.webscraping/webscraping-ai
import ai.webscraping.Client;
import ai.webscraping.Config;
import ai.webscraping.option.TextOptions;
Client client = new Client(Config.builder().apiKey("YOUR_API_KEY").build());
String text = client.text(TextOptions.builder()
.url("https://example.com")
.textFormat("json")
.build());
System.out.println(text);// dotnet add package WebScrapingAI
// https://www.nuget.org/packages/WebScrapingAI
using WebScrapingAI;
var client = new WebScrapingAIClient(new WebScrapingAIClientOptions { ApiKey = "YOUR_API_KEY" });
var text = await client.TextAsync(new TextRequest {
Url = "https://example.com",
TextFormat = "json",
});
Console.WriteLine(text);Extract HTML from specific page elements using CSS selectors. Useful when you only need a portion of the page.
/selected
| Parameter | Type | Description |
|---|---|---|
| url required | string | URL of the webpage |
| selector required | string | CSS selector (e.g., "h1", ".price", "#main") |
curl -G "https://api.webscraping.ai/selected" \
--data-urlencode "api_key=YOUR_API_KEY" \
--data-urlencode "url=https://example.com" \
--data-urlencode "selector=h1"# pip install webscraping_ai
# https://pypi.org/project/webscraping-ai/
from webscraping_ai import Client
client = Client(api_key="YOUR_API_KEY")
html = client.selected("https://example.com", selector="h1")
print(html)// npm install webscraping-ai
// https://www.npmjs.com/package/webscraping-ai
import { WebScrapingAI } from 'webscraping-ai';
const client = new WebScrapingAI({ apiKey: 'YOUR_API_KEY' });
const html = await client.selected({
url: 'https://example.com',
selector: 'h1',
});
console.log(html);<?php
// composer require webscraping-ai/webscraping-ai-php
// https://packagist.org/packages/webscraping-ai/webscraping-ai-php
require 'vendor/autoload.php';
use WebScrapingAI\Client;
$client = new Client('YOUR_API_KEY');
$html = $client->selected('https://example.com', 'h1');
echo $html;# gem install webscraping_ai
# https://rubygems.org/gems/webscraping_ai
require 'webscraping_ai'
client = WebScrapingAI::Client.new(api_key: 'YOUR_API_KEY')
html = client.selected('https://example.com', selector: 'h1')
puts html// go get github.com/webscraping-ai/webscraping-ai-go/v4
// https://pkg.go.dev/github.com/webscraping-ai/webscraping-ai-go/v4
package main
import (
"context"
"fmt"
webscrapingai "github.com/webscraping-ai/webscraping-ai-go/v4"
)
func main() {
client, _ := webscrapingai.NewClient(&webscrapingai.Config{APIKey: "YOUR_API_KEY"})
html, _ := client.Selected(context.Background(), &webscrapingai.SelectedOptions{
URL: "https://example.com",
Selector: "h1",
})
fmt.Println(html)
}// Maven: ai.webscraping:webscraping-ai:4.2.0
// https://central.sonatype.com/artifact/ai.webscraping/webscraping-ai
import ai.webscraping.Client;
import ai.webscraping.Config;
import ai.webscraping.option.SelectedOptions;
Client client = new Client(Config.builder().apiKey("YOUR_API_KEY").build());
String html = client.selected(SelectedOptions.builder()
.url("https://example.com")
.selector("h1")
.build());
System.out.println(html);// dotnet add package WebScrapingAI
// https://www.nuget.org/packages/WebScrapingAI
using WebScrapingAI;
var client = new WebScrapingAIClient(new WebScrapingAIClientOptions { ApiKey = "YOUR_API_KEY" });
var html = await client.SelectedAsync(new SelectedRequest {
Url = "https://example.com",
Selector = "h1",
});
Console.WriteLine(html);/selected-multiple endpoint to extract several page areas in one request.
Extract several page areas in one request: pass one CSS selector per area and get back, for each selector, the HTML of every element it matches. Costs the same as a single /selected request.
/selected-multiple
| Parameter | Type | Description |
|---|---|---|
| url required | string | URL of the webpage |
| selectors required | string[] | CSS selectors. Repeat the parameter once per selector (?selectors=h1&selectors=p), not selectors[]=, which is ignored. |
curl -G "https://api.webscraping.ai/selected-multiple" \
--data-urlencode "api_key=YOUR_API_KEY" \
--data-urlencode "url=https://example.com" \
--data-urlencode "selectors=h1" \
--data-urlencode "selectors=p"# pip install webscraping_ai
# https://pypi.org/project/webscraping-ai/
from webscraping_ai import Client
client = Client(api_key="YOUR_API_KEY")
areas = client.selected_multiple("https://example.com", selectors=["h1", "p"])
print(areas)// npm install webscraping-ai
// https://www.npmjs.com/package/webscraping-ai
import { WebScrapingAI } from 'webscraping-ai';
const client = new WebScrapingAI({ apiKey: 'YOUR_API_KEY' });
const areas = await client.selectedMultiple({
url: 'https://example.com',
selectors: ['h1', 'p'],
});
console.log(areas);<?php
// composer require webscraping-ai/webscraping-ai-php
// https://packagist.org/packages/webscraping-ai/webscraping-ai-php
require 'vendor/autoload.php';
use WebScrapingAI\Client;
$client = new Client('YOUR_API_KEY');
$areas = $client->selectedMultiple('https://example.com', ['h1', 'p']);
print_r($areas);# gem install webscraping_ai
# https://rubygems.org/gems/webscraping_ai
require 'webscraping_ai'
client = WebScrapingAI::Client.new(api_key: 'YOUR_API_KEY')
areas = client.selected_multiple('https://example.com', selectors: ['h1', 'p'])
p areas// go get github.com/webscraping-ai/webscraping-ai-go/v4
// https://pkg.go.dev/github.com/webscraping-ai/webscraping-ai-go/v4
package main
import (
"context"
"fmt"
webscrapingai "github.com/webscraping-ai/webscraping-ai-go/v4"
)
func main() {
client, _ := webscrapingai.NewClient(&webscrapingai.Config{APIKey: "YOUR_API_KEY"})
areas, _ := client.SelectedMultiple(context.Background(), &webscrapingai.SelectedMultipleOptions{
URL: "https://example.com",
Selectors: []string{"h1", "p"},
})
fmt.Println(areas)
}// Maven: ai.webscraping:webscraping-ai:4.2.0
// https://central.sonatype.com/artifact/ai.webscraping/webscraping-ai
import ai.webscraping.Client;
import ai.webscraping.Config;
import ai.webscraping.option.SelectedMultipleOptions;
import ai.webscraping.result.SelectedMultipleResult;
Client client = new Client(Config.builder().apiKey("YOUR_API_KEY").build());
SelectedMultipleResult areas = client.selectedMultiple(SelectedMultipleOptions.builder()
.url("https://example.com")
.selectors("h1", "p")
.build());
System.out.println(areas.getResults());// dotnet add package WebScrapingAI
// https://www.nuget.org/packages/WebScrapingAI
using WebScrapingAI;
var client = new WebScrapingAIClient(new WebScrapingAIClientOptions { ApiKey = "YOUR_API_KEY" });
var areas = await client.SelectedMultipleAsync(new SelectedMultipleRequest {
Url = "https://example.com",
Selectors = new[] { "h1", "p" },
});
foreach (var matches in areas.Results) Console.WriteLine(string.Join(" | ", matches));One array per selector, in request order, holding the inner HTML of every element that matches it. A selector that matches nothing returns an empty array.
Get parsed Google search results for a query as JSON: organic results (position, title, link, domain, displayed link, snippet, date), related searches, spelling-correction info, and pagination. Google is the only engine today.
Unlike the scraping endpoints, /serp is query-shaped, not URL-shaped: you pass the search query in q instead of a url, and WebScraping.AI handles proxy routing and parsing. The page-scraping parameters (js, proxy, country, headers, timeout, …) don't apply. Each successful search costs a flat 15 credits; failed searches are not charged.
/serp
| Parameter | Type | Description |
|---|---|---|
| q required | string | Search query, e.g. coffee machines |
| api_key required | string | Your API key |
| engine | string | Search engine to query. Only google (the default) is supported |
| gl | string | Two-letter country code for the search location (Google gl). Default: us |
| hl | string | Two-letter language code for the results (Google hl). Default: en |
| page | integer | Results page number, 10 results per page. A whole number from 1 to 100. Default: 1. Any other value (0, a fraction, a non-number, or above 100) is rejected with a 400 before billing |
curl -G "https://api.webscraping.ai/serp" \
--data-urlencode "api_key=YOUR_API_KEY" \
--data-urlencode "q=coffee machines" \
--data-urlencode "gl=us" \
--data-urlencode "hl=en" \
--data-urlencode "page=1"# pip install webscraping_ai
# https://pypi.org/project/webscraping-ai/
from webscraping_ai import Client
client = Client(api_key="YOUR_API_KEY")
results = client.serp("coffee machines", gl="us", hl="en", page=1)
for r in results["organic_results"]:
print(r["position"], r["title"], r["link"])// npm install webscraping-ai
// https://www.npmjs.com/package/webscraping-ai
import { WebScrapingAI } from 'webscraping-ai';
const client = new WebScrapingAI({ apiKey: 'YOUR_API_KEY' });
const serp = await client.serp({ q: 'coffee machines', gl: 'us', hl: 'en', page: 1 });
for (const r of serp.organic_results) {
console.log(`${r.position}. ${r.title} — ${r.link}`);
}<?php
// composer require webscraping-ai/webscraping-ai-php
// https://packagist.org/packages/webscraping-ai/webscraping-ai-php
require 'vendor/autoload.php';
use WebScrapingAI\Client;
$client = new Client('YOUR_API_KEY');
$serp = $client->serp(q: 'coffee machines', gl: 'us', hl: 'en', page: 1);
foreach ($serp['organic_results'] as $result) {
printf("%d. %s — %s\n", $result['position'], $result['title'], $result['link']);
}# gem install webscraping_ai
# https://rubygems.org/gems/webscraping_ai
require 'webscraping_ai'
client = WebScrapingAI::Client.new(api_key: 'YOUR_API_KEY')
results = client.serp(q: 'coffee machines', gl: 'us', hl: 'en', page: 1)
results['organic_results'].each do |r|
puts "#{r['position']}. #{r['title']} — #{r['link']}"
end// go get github.com/webscraping-ai/webscraping-ai-go/v4
// https://pkg.go.dev/github.com/webscraping-ai/webscraping-ai-go/v4
package main
import (
"context"
"fmt"
webscrapingai "github.com/webscraping-ai/webscraping-ai-go/v4"
)
func main() {
client, _ := webscrapingai.NewClient(&webscrapingai.Config{APIKey: "YOUR_API_KEY"})
page := 1
serp, _ := client.Serp(context.Background(), &webscrapingai.SerpOptions{
Q: "coffee machines",
GL: "us",
HL: "en",
Page: &page,
})
for _, r := range serp.OrganicResults {
fmt.Println(r.Position, r.Title, r.Link)
}
}// Maven: ai.webscraping:webscraping-ai:4.2.0
// https://central.sonatype.com/artifact/ai.webscraping/webscraping-ai
import ai.webscraping.Client;
import ai.webscraping.Config;
import ai.webscraping.option.SerpOptions;
import ai.webscraping.result.SerpResult;
Client client = new Client(Config.builder().apiKey("YOUR_API_KEY").build());
SerpResult serp = client.serp(SerpOptions.builder()
.q("coffee machines")
.gl("us")
.hl("en")
.page(1)
.build());
for (SerpResult.OrganicResult r : serp.getOrganicResults()) {
System.out.println(r.getPosition() + ". " + r.getTitle() + " — " + r.getLink());
}// dotnet add package WebScrapingAI
// https://www.nuget.org/packages/WebScrapingAI
using WebScrapingAI;
var client = new WebScrapingAIClient(new WebScrapingAIClientOptions { ApiKey = "YOUR_API_KEY" });
var serp = await client.SerpAsync(new SerpRequest {
Q = "coffee machines",
Gl = "us",
Hl = "en",
Page = 1,
});
foreach (var r in serp.OrganicResults) {
Console.WriteLine($"{r.Position}. {r.Title} — {r.Link}");
}Trimmed to two results; a full page has up to 10. snippet, date, related_searches and search_information.showing_results_for (the auto-corrected query) appear only when Google shows them, and pagination.next is absent on the last page. People Also Ask, ads, knowledge graph and local results are not included.
position restarts at 1 on every page. For an absolute rank, compute (page - 1) * 10 + position. A search that genuinely returns no results (organic_results_state is Fully empty) is a successful, billed search; only failed searches are free.
Get structured JSON for a public page on a supported site from its normal URL: for example a YouTube video, channel or playlist, a TikTok video or profile, an X post or profile, a LinkedIn company, job or profile, an Instagram post, reel or profile, or a Reddit post, subreddit or user. The site and page type are detected from the URL, and WebScraping.AI handles fetching, proxies and parsing.
More sites are added over time, so don't validate URLs on your side: an unsupported URL or page type returns a 400 that is not charged, whose message lists what is supported. For other sites, use /ai/fields. The page-scraping parameters (js, proxy, headers, timeout, …) don't apply. Each request costs 15 credits (50 for Reddit, whose pages are loaded in a real browser); requests that fail to fetch are not charged, while pages that parse empty or don't exist are.
Site guides (supported page types, fields and limits per site): Social Media Scraper API (overview of all sites), YouTube Scraper API, TikTok Scraper API, Twitter (X) Scraper API, LinkedIn Scraper API, Instagram Scraper API, Reddit Scraper API.
/data
| Parameter | Type | Description |
|---|---|---|
| url required | string | URL of a page on a supported site, e.g. https://www.youtube.com/watch?v=dQw4w9WgXcQ |
| api_key required | string | Your API key |
| country | string | Country of the proxy used to fetch the page, one of the proxy countries. Default: us |
| transcript | boolean | YouTube videos only. Also fetch the transcript into data.transcript (null when no matching captions are available). If the transcript fetch fails, the whole request fails with a 500 and is not charged |
| transcript_language | string | YouTube videos only, with transcript=true. Caption language to pick, e.g. de. Without it, English is preferred, then the first available track |
curl -G "https://api.webscraping.ai/data" \
--data-urlencode "api_key=YOUR_API_KEY" \
--data-urlencode "url=https://www.youtube.com/watch?v=dQw4w9WgXcQ"# pip install webscraping_ai
# https://pypi.org/project/webscraping-ai/
from webscraping_ai import Client
client = Client(api_key="YOUR_API_KEY")
result = client.data("https://www.youtube.com/watch?v=dQw4w9WgXcQ")
print(result["request_parameters"]["provider"], result["parse_status"])
if result["data"]:
print(result["data"]["title"])// npm install webscraping-ai
// https://www.npmjs.com/package/webscraping-ai
import { WebScrapingAI } from 'webscraping-ai';
const client = new WebScrapingAI({ apiKey: 'YOUR_API_KEY' });
const result = await client.data({ url: 'https://www.youtube.com/watch?v=dQw4w9WgXcQ' });
console.log(result.request_parameters.provider, result.parse_status);
console.log(result.data?.title);<?php
// composer require webscraping-ai/webscraping-ai-php
// https://packagist.org/packages/webscraping-ai/webscraping-ai-php
require 'vendor/autoload.php';
use WebScrapingAI\Client;
$client = new Client('YOUR_API_KEY');
$result = $client->data(url: 'https://www.youtube.com/watch?v=dQw4w9WgXcQ');
echo $result['request_parameters']['provider'], ' ', $result['parse_status'], "\n";
echo $result['data']['title'] ?? '', "\n";# gem install webscraping_ai
# https://rubygems.org/gems/webscraping_ai
require 'webscraping_ai'
client = WebScrapingAI::Client.new(api_key: 'YOUR_API_KEY')
result = client.data('https://www.youtube.com/watch?v=dQw4w9WgXcQ')
puts "#{result['request_parameters']['provider']} #{result['parse_status']}"
puts result['data']&.dig('title')// go get github.com/webscraping-ai/webscraping-ai-go/v4
// https://pkg.go.dev/github.com/webscraping-ai/webscraping-ai-go/v4
package main
import (
"context"
"encoding/json"
"fmt"
webscrapingai "github.com/webscraping-ai/webscraping-ai-go/v4"
)
func main() {
client, _ := webscrapingai.NewClient(&webscrapingai.Config{APIKey: "YOUR_API_KEY"})
res, _ := client.Data(context.Background(), &webscrapingai.DataOptions{
URL: "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
})
fmt.Println(res.RequestParameters.Provider, res.ParseStatus)
// Data is raw JSON whose shape depends on the provider and page type
var video struct {
Title string `json:"title"`
}
if res.Data != nil {
json.Unmarshal(res.Data, &video)
}
fmt.Println(video.Title)
}// Maven: ai.webscraping:webscraping-ai:4.2.0
// https://central.sonatype.com/artifact/ai.webscraping/webscraping-ai
import ai.webscraping.Client;
import ai.webscraping.Config;
import ai.webscraping.option.DataOptions;
import ai.webscraping.result.DataResult;
Client client = new Client(Config.builder().apiKey("YOUR_API_KEY").build());
DataResult result = client.data(DataOptions.builder()
.url("https://www.youtube.com/watch?v=dQw4w9WgXcQ")
.build());
System.out.println(result.getRequestParameters().getProvider() + " " + result.getParseStatus());
if (result.getData() != null) {
System.out.println(result.getData().path("title").asText());
}// dotnet add package WebScrapingAI
// https://www.nuget.org/packages/WebScrapingAI
using WebScrapingAI;
var client = new WebScrapingAIClient(new WebScrapingAIClientOptions { ApiKey = "YOUR_API_KEY" });
var result = await client.DataAsync(new DataRequest {
Url = "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
});
Console.WriteLine($"{result.RequestParameters.Provider} {result.ParseStatus}");
if (result.Data is { } data) {
Console.WriteLine(data.GetProperty("title").GetString());
}Trimmed. The fields in data depend on provider and type; they use snake_case, and fields a page doesn't expose are null (some flags come back false, and lists come back empty).
ok when the page was parsed, parse_failed when it was fetched but couldn't be parsed (data may be null or partial), not_found when the page doesn't exist. All three are successful, billed requests; only failed fetches are free.
Check your remaining API credits and account status.
/account
curl "https://api.webscraping.ai/account?api_key=YOUR_API_KEY"# pip install webscraping_ai
# https://pypi.org/project/webscraping-ai/
from webscraping_ai import Client
client = Client(api_key="YOUR_API_KEY")
info = client.account()
print(info)// npm install webscraping-ai
// https://www.npmjs.com/package/webscraping-ai
import { WebScrapingAI } from 'webscraping-ai';
const client = new WebScrapingAI({ apiKey: 'YOUR_API_KEY' });
const info = await client.account();
console.log(info);<?php
// composer require webscraping-ai/webscraping-ai-php
// https://packagist.org/packages/webscraping-ai/webscraping-ai-php
require 'vendor/autoload.php';
use WebScrapingAI\Client;
$client = new Client('YOUR_API_KEY');
$info = $client->account();
print_r($info);# gem install webscraping_ai
# https://rubygems.org/gems/webscraping_ai
require 'webscraping_ai'
client = WebScrapingAI::Client.new(api_key: 'YOUR_API_KEY')
info = client.account
puts info.inspect// go get github.com/webscraping-ai/webscraping-ai-go/v4
// https://pkg.go.dev/github.com/webscraping-ai/webscraping-ai-go/v4
package main
import (
"context"
"fmt"
webscrapingai "github.com/webscraping-ai/webscraping-ai-go/v4"
)
func main() {
client, _ := webscrapingai.NewClient(&webscrapingai.Config{APIKey: "YOUR_API_KEY"})
info, _ := client.Account(context.Background())
fmt.Printf("%s - %d calls left\n", info.Email, info.RemainingAPICalls)
}// Maven: ai.webscraping:webscraping-ai:4.2.0
// https://central.sonatype.com/artifact/ai.webscraping/webscraping-ai
import ai.webscraping.Client;
import ai.webscraping.Config;
import ai.webscraping.result.AccountInfo;
Client client = new Client(Config.builder().apiKey("YOUR_API_KEY").build());
AccountInfo info = client.account();
System.out.println(info);// dotnet add package WebScrapingAI
// https://www.nuget.org/packages/WebScrapingAI
using WebScrapingAI;
var client = new WebScrapingAIClient(new WebScrapingAIClientOptions { ApiKey = "YOUR_API_KEY" });
var info = await client.AccountAsync();
Console.WriteLine($"{info.Email} - {info.RemainingApiCalls} calls left");Understanding API error responses and how to handle them.
| Code | Description | Solution |
|---|---|---|
| 400 | Invalid parameters | Check parameter values and format |
| 402 | Insufficient credits | Upgrade your plan, top up pay-as-you-go credits, or wait for credit reset |
| 403 | Invalid API key | Verify your API key is correct |
| 429 | Too many concurrent requests | Reduce request rate or upgrade for higher concurrency |
| 500 | Target page could not be scraped (blocked, anti-bot challenge, non-2xx status, DNS failure), or an unexpected error (error_code internal_error) |
Read error_code and follow next_step (see below); retry an internal_error and contact support if it persists |
| 504 | Request timeout | Increase timeout parameter value |
When the target website blocks the request or returns an error, the 500 response body tells you what happened and what to try next, with the real credit cost of that retry. Example for a site that blocks datacenter IPs:
{
"message": "The target website returned HTTP 403 (Forbidden) through a datacenter proxy: it is blocking this IP pool (anti-bot protection). Retry with proxy=residential (10 credits per request). Failed requests are not billed.",
"error_code": "target_blocked",
"status_code": 403,
"status_message": "Forbidden",
"body": "<html><body>Access denied</body></html>",
"request_parameters": { "proxy": "datacenter", "js": false },
"next_step": {
"params": { "proxy": "residential" },
"cost": 10,
"message": "Retry with proxy=residential (10 credits per request)."
}
}
The escalation ladder is datacenter → residential → stealth. Anti-bot challenge pages go straight to proxy=stealth, because they need the hardened browser, not just a better IP. When the request already ran on the stealth tier (or a dedicated route), next_step is omitted and the message says the target is currently not reachable; Google search URLs are pointed at the /serp endpoint instead. Automating the retry is safe: next_step.params are ordinary request parameters to set, next_step.remove (present only when you used custom_proxy, which overrides any proxy setting) lists parameters to drop, and next_step.cost is the price per request on current plans if the retry succeeds. If you would rather not automate it yourself, proxy=auto walks the same ladder inside one request (see the proxy parameter); its error bodies carry request_parameters.auto: true and request_parameters.tiers_tried, and request_parameters.proxy is the tier the walk ended on.
error_code |
Meaning | Next step |
|---|---|---|
target_blocked | Target returned 403, 429, 503 or 999: an IP-reputation block or rate limit | Better proxy tier from next_step |
target_challenge | Target served an anti-bot challenge page (Cloudflare, Akamai, PerimeterX, Datadome) instead of content, or closed the browser tab | proxy=stealth; a Cloudflare JavaScript challenge hit with js=false suggests js=true instead (with proxy=residential on datacenter) |
target_error | Other non-2xx status from the target (its own 4xx or 5xx) | Check the URL; for 5xx a proxy upgrade is suggested as a hedge |
target_redirect | Redirect returned because error_on_redirect=true | Billed as a success; the redirect target is in message |
target_unreachable | Hostname does not resolve, or the target dropped the connection without an HTTP response (range-level blocking) | Check the URL, or the proxy tier from next_step |
wait_for_timeout | wait_for selector never appeared | Check the selector or increase timeout |
invalid_url, forbidden_target, custom_proxy_error | Malformed URL, a domain/body the API does not serve, or your custom_proxy rejected our credentials (407) | Fix the request |
response_too_large | Response over the size limit | Use /selected or /selected-multiple |
timeout (504), internal_error | Page did not load in time, or an error on our side | Increase timeout; for internal errors retry or contact support |
Complete documentation for all API parameters of the scraping and AI endpoints. Click on any parameter to see detailed information and examples. The search results endpoint takes its own query parameters (q, engine, gl, hl, page), documented in its section, and so does the structured data endpoint (url, country, transcript, transcript_language).
| Parameter | Type | Default | Description |
|---|---|---|---|
| url | string | required | Target webpage URL |
| api_key | string | required | Your API key for authentication |
| js | boolean | true | Enable JavaScript rendering |
| js_timeout | integer | 2000 | JavaScript rendering timeout (ms) |
| timeout | integer | 10000 | Total request timeout (ms) |
| wait_for | string | - | CSS selector to wait for |
| proxy | string | datacenter | Proxy type |
| country | string | us | Proxy country |
| device | string | desktop | Device emulation |
| headers | object | - | Custom HTTP headers |
| js_script | string | - | Custom JavaScript to execute |
| custom_proxy | string | - | Your own proxy URL |
| error_on_404 | boolean | false | Return error for 404 pages |
| error_on_redirect | boolean | false | Return error on redirects |
The URL of the target webpage to scrape or analyze. Must be a valid HTTP or HTTPS URL.
Detailsurl=https://example.com/page?id=123
Your unique API key for authentication. Get your key from the dashboard.
Enable JavaScript rendering using a headless Chromium browser. Required for SPAs, dynamic content, and modern web applications.
When to use js=true# With JS rendering (default)
curl "https://api.webscraping.ai/html?api_key=KEY&url=https://spa-app.com&js=true"
# Without JS rendering (faster, cheaper)
curl "https://api.webscraping.ai/html?api_key=KEY&url=https://static-site.com&js=false"
Maximum time in milliseconds to wait for JavaScript execution after the page loads. Increase this value if you see loading indicators instead of actual content.
Detailsjs=truewait_for for more precisionMaximum total time in milliseconds for the entire request, including page retrieval, JavaScript rendering, and processing.
Detailstimeout=15000
Sets a 15-second timeout for slow-loading pages
CSS selector to wait for before returning the page content. The request will wait until this element appears in the DOM, then return the content. This overrides js_timeout.
If the element never appears within the request timeout, the request fails with a 500 error naming the selector, along with the target page's HTTP status code (status_code) and a preview of the page body (body) so you can see what the page actually contained. Failed requests are not billed. If you see this error on a page that should contain the element, check the selector, increase the timeout parameter, or inspect the body preview for anti-bot challenges.
wait_for=.product-pricewait_for=#search-resultswait_for=[data-loaded="true"]wait_for=.reviews-container# Wait for product grid to load before scraping
curl -G "https://api.webscraping.ai/html" \
--data-urlencode "api_key=YOUR_API_KEY" \
--data-urlencode "url=https://shop.com/products" \
--data-urlencode "wait_for=.product-grid"
Type of proxy to use for the request. Choose based on your target website's anti-bot measures.
| Value | Description | Best for | Cost |
|---|---|---|---|
datacenter |
Fast datacenter proxies with rotating IPs | Most websites, APIs, general scraping | 1 credit (5 with JS) |
residential |
Real residential IPs from ISPs | Anti-bot protected sites, sneaker sites, social media | 10 credits (25 with JS) |
stealth |
Premium proxies with advanced anti-bot bypass | The most heavily protected sites where residential is not enough | 50 credits |
auto |
Tries datacenter, then residential, then stealth inside one request and returns the first result that isn't blocked | Unknown or mixed targets, first runs against a new site | The tier that worked (1–50 credits); failed attempts are free |
proxy=auto does this walk for you: anti-bot challenge pages skip straight to stealth, and the API remembers the cheapest tier that worked for each domain for 30 days, so later requests to a protected site start there instead of re-trying the blocked tiers. Each extra tier costs response time, so once you know a site needs residential, setting it explicitly is faster. auto cannot be combined with custom_proxy.
Country code for geo-targeting. The request will be made from a proxy in the specified country.
| Code | Country | Code | Country |
|---|---|---|---|
us |
United States | ru |
Russia |
gb |
United Kingdom | jp |
Japan |
de |
Germany | kr |
South Korea |
fr |
France | in |
India |
ca |
Canada | it |
Italy |
es |
Spain | hk |
Hong Kong |
tr |
Turkey |
# Scrape from a UK IP address
curl "https://api.webscraping.ai/html?api_key=KEY&url=https://uk-shop.com&country=gb"
Device type emulation. Affects viewport size, user agent, and touch capabilities.
| Value | Viewport | Use case |
|---|---|---|
desktop |
1920x940 (1080p screen) | Desktop websites, full layouts |
mobile |
414x715 (iPhone 11) | Mobile sites, responsive layouts, AMP pages |
tablet |
810x1080 (iPad 7th gen) | Tablet-optimized layouts |
Custom HTTP headers to send with the request. Useful for authentication, cookies, or custom user agents.
Format optionsheaders={"Cookie":"session=abc"}headers[Cookie]=session=abcCookie - Session cookiesAuthorization - Bearer tokensReferer - Referrer URL# Pass custom headers as JSON
curl -G "https://api.webscraping.ai/html" \
--data-urlencode "api_key=YOUR_API_KEY" \
--data-urlencode "url=https://example.com" \
--data-urlencode 'headers={"Cookie":"session=abc123","Authorization":"Bearer token"}'
Custom JavaScript code to execute on the page after it loads. Useful for clicking buttons, scrolling, filling forms, or extracting data. Only the /html endpoint runs it; the other endpoints ignore it. The script's result is the value of its last expression, so don't use a top-level return (it is a syntax error).
return_script_result=true to get the return value of your script instead of the page HTML.
# Click a button
js_script=document.querySelector('.load-more-btn').click()
# Scroll to bottom
js_script=window.scrollTo(0, document.body.scrollHeight)
# Extract data and return it (with return_script_result=true)
js_script=JSON.stringify(window.__INITIAL_DATA__)
Use your own proxy server instead of our built-in proxy pool. Useful if you have specific proxy requirements or existing proxy subscriptions.
Formathttp://username:password@host:port
Example
custom_proxy=http://user:pass@proxy.example.com:8080
Return an error response when the target page returns a 404 status code, instead of returning the 404 page content.
Return an error when the target page redirects, instead of following the redirect.
For the complete machine-readable API specification — every endpoint, parameter, and response schema — download our OpenAPI 3.1 spec.
Download OpenAPI Spec