Web Scraping Glossary
Plain-English definitions for the scraping, proxy, and anti-bot terms you'll run into building a scraper.
43 terms and counting
A
-
AI Crawler
Learn the definition of an AI crawler, an automated bot operated by an AI company that visits websites to collect content for training large language models, powering real-time answers, or generating citations in AI search products.
Read more → -
Anti-Bot System
An anti-bot system is a layered defense mechanism analyzing IP reputation, fingerprints, and behavior to block automated scrapers.
Read more → -
Anti-Detect Browser
An anti-detect browser isolates multi-accounting profiles with spoofed fingerprints. Compare anti-detect browsers with automated scraping browsers.
Read more → -
Anti-Fingerprinting
Learn the definition of anti-fingerprinting, spoofing, or randomizing canvas, WebGL, fonts, and screen attributes so a browser doesn't produce the same identifiable hash.
Read more → -
ASN (Autonomous System Number)
An ASN identifies the network operator owning an IP block, allowing anti-bot systems to classify traffic as residential ISP or datacenter.
Read more →
B
-
Behavioral Biometrics
Behavioral biometrics analyzes mouse movements and keystroke dynamics to detect bots. Learn how anti-bot systems use behavioral signals to flag scrapers.
Read more → -
Browser Fingerprinting
Browser fingerprinting builds a composite identity from hardware and rendering signals to detect scrapers, persisting even when IP addresses change.
Read more →
C
-
Canvas Fingerprinting
Canvas fingerprinting extracts pixel rendering quirks from HTML5 canvases to identify scrapers across sessions. Learn detection and spoofing methods.
Read more → -
CGNAT
CGNAT enables mobile carriers to share public IPs across thousands of devices, explaining why mobile proxy pools resist anti-bot blacklists and blocks.
Read more → -
Crawl Depth
Learn what crawl depth means and how to set the max crawl depth for both SEO audits and web scraping jobs.
Read more → -
CSS Selector
A CSS selector targets HTML elements by class, ID, or attributes in web scraping. Compare CSS selector vs XPath performance and essential scraper syntax.
Read more →
D
-
Datacenter Proxy
Learn the definition of a datacenter proxy, an IP hosted in a commercial data center rather than an ISP, and how it compares to residential proxies for web scraping.
Read more → -
DNS Leak
A DNS leak occurs when domain lookups bypass a proxy tunnel, exposing real IP addresses. Learn how DNS leak tests and socks5h resolve scraper leaks.
Read more →
E
-
Elite Proxy (High-Anonymity Proxy)
Learn the definition of an elite proxy, a high-anonymity proxy that hides your real IP and strips the headers that reveal a proxy is being used at all.
Read more →
G
-
Good Bot vs Bad Bot
Learn the difference between good bot vs bad bot traffic, how modern firewalls classify automated crawlers, and how scrapers navigate anti-bot detection.
Read more →
H
-
Headless Browser
Learn the definition of a headless browser, a full browser engine that renders pages and runs JavaScript without a visible UI and why anti-bot systems can fingerprint it.
Read more → -
Headless Browser Detection
Learn how sites detect headless browsers, navigator.webdriver, empty plugin arrays, WebGL renderer strings, and the JavaScript probes that identify automation.
Read more → -
Honeypot Trap
Learn the definition of a honeypot trap, a link or form field hidden from human users that flags automated scrapers and bots the moment they interact with it.
Read more →
I
-
Infinite Scroll Scraping
Learn how infinite scroll scraping works, why a plain HTTP request only ever sees the first batch of content, and how to script scrolling in a headless browser.
Read more → -
IP Blacklist (IP Blocklist)
An IP blacklist is a database of flagged IP addresses blocked by firewalls. Learn how to check if an IP is blacklisted and prevent scraping bans.
Read more → -
IP Reputation
IP reputation is a trust score assigned to IP addresses based on ASN type and abuse history, determining whether scrapers get challenged or allowed.
Read more → -
IP Rotation
IP rotation cycles requests through a large proxy pool to avoid rate limits and bans. Understand how rotating IP addresses work for web scraping.
Read more →
J
-
JavaScript Rendering
JavaScript rendering builds page content dynamically in the browser, requiring headless browsers to extract data from modern client-side apps.
Read more →
L
-
llms.txt
Learn the definition of llms.txt, a proposed markdown file format placed at a website's root to provide AI models and large language models with a curated summary of key content and documentation.
Read more →
P
-
Pagination (Web Scraping Pagination)
Learn what pagination means in web scraping, how numbered, next-link, and cursor patterns differ, and how to work through paged results without missing records.
Read more → -
Proxy Authentication
Proxy authentication verifies client access via user:password credentials or IP whitelisting before routing requests to prevent 407 unauthorized proxy errors.
Read more → -
Proxy Bandwidth
Proxy bandwidth is the data volume in GB metered by residential proxy providers. Learn how page weight and media assets drive proxy bandwidth cost.
Read more → -
Proxy Chaining
Proxy chaining routes network traffic through sequential intermediary proxies to hide origin IPs, balancing nested anonymity against connection latency.
Read more → -
Proxy Pool
A proxy pool is the set of IP addresses a scraper rotates across requests to distribute traffic and avoid blocks.
Read more → -
Public Data Scraping
Learn what public data scraping means, how hiQ v. LinkedIn shaped the legal picture, and why public access does not mean unrestricted use of the data.
Read more →
R
-
Reverse Proxy
Learn the definition of a reverse proxy, a server that sits in front of backend servers and forwards client requests to them, and how it differs from a forward proxy.
Read more → -
Robots Meta Tag (noindex/nofollow)
Learn what a robots meta tag is, how noindex and nofollow work, and why it differs from robots.txt, which blocks crawling before a page is ever fetched.
Read more →
S
-
Selectorless Scraping (AI Web Scraping)
Learn what selectorless scraping is, how AI web scraping extracts data from a plain-language description instead of CSS or XPath selectors, and its trade-offs.
Read more → -
Stealth Mode
Learn the definition of stealth mode in browser automation, the patches that hide navigator.webdriver and other automation signals, and what they can't fix.
Read more → -
Sticky Session (Session Persistence)
Technical definition of sticky sessions in web scraping: IP persistence mechanics, TTL specifications, and session ID proxy parameters.
Read more →
T
-
Terms of Service (ToS) Violation
Learn what a Terms of Service violation means for web scraping, why it is a contract question rather than a criminal one, and how it differs from copyright claims.
Read more → -
TLS Fingerprinting (JA3 / JA4)
Learn the definition of TLS fingerprinting, a security mechanism that analyzes initial handshake patterns to identify and block automated clients using standards like JA3 and JA4, regardless of IP rotation or headers.
Read more → -
Transparent Proxy
Learn the definition of a transparent proxy, an intercepting proxy that forwards your real IP and identifies itself, offering no anonymity for web scraping.
Read more →
U
-
User Agent
Technical definition of User-Agent headers in web scraping: string format anatomy, modern UA reference list, Client Hints (Sec-CH-UA), and browser fingerprint alignment.
Read more →
V
-
VPN vs Proxy
VPN vs proxy explained: a VPN encrypts all device traffic through one tunnel, while a proxy routes single-app traffic and scales across rotating IPs.
Read more →
W
-
Web Crawler
A web crawler systematically discovers and follows links across the web to build an index, differentiating it from a web scraper.
Read more → -
Web Scraping Pipeline
Learn what a web scraping pipeline is, the stages a production scraper runs through from fetch to storage, and why separating them keeps a scraper maintainable.
Read more →
X
-
XPath
Learn what XPath is, how it selects nodes by structure and content in HTML, and how it compares to CSS selectors for web scraping, with a syntax cheat sheet.
Read more →
No terms match that search yet.
More terms are added every week.