Public Data Scraping
Public data scraping is collecting information that's visible without logging in or bypassing any access control. Courts in cases like hiQ v. LinkedIn have generally treated this differently from scraping content behind authentication, though the legal picture keeps evolving, as the ongoing Reddit v. Oxylabs case over scraping public data shows.
This page is general information, not legal advice. The law here is unsettled and jurisdiction-specific. Talk to a lawyer about your specific situation.
What Counts as Public Data?
The working test is whether a request without credentials returns the content. What is public data in practice covers product listings and prices, published articles, company pages, job postings, public social profiles, government records, and search results. What it excludes is anything behind a login, a paywall, an API key, or a permission check.
The distinction is about access, not sensitivity. A public LinkedIn profile is public data even though it's about a person. An internal price list is not public data even though it contains no personal information at all.
hiQ v. LinkedIn and the Access Question
hiQ Labs scraped public LinkedIn profiles; LinkedIn sent a cease-and-desist and blocked it. The Ninth Circuit held that scraping publicly available data likely doesn't constitute access "without authorization" under the Computer Fraud and Abuse Act, reasoning that data open to anyone with a browser isn't gated in the way the statute contemplates. The Supreme Court's 2021 decision in Van Buren, which read the CFAA's authorization language narrowly, pointed in a compatible direction.
hiQ v. LinkedIn is frequently cited as establishing that public scraping is legal. The ending complicates that. On remand, the case turned to LinkedIn's user agreement rather than the CFAA, hiQ was found to have breached it, and the parties settled in late 2022 with hiQ agreeing to an injunction.
So the case supports a narrow proposition: scraping public data is unlikely to be a computer-fraud problem. It says little about contract claims, which is exactly where hiQ eventually lost.
Reddit v. Oxylabs and the Current Direction
In October 2025, Reddit filed suit in the Southern District of New York against Perplexity AI, Oxylabs, SerpApi, and AWMProxy, alleging industrial-scale scraping of Reddit content collected from Google search results rather than from Reddit directly. Oxylabs has publicly contested the claims, arguing that no company owns publicly available data.
What makes the case worth watching is its legal theory. Reddit leans on the DMCA's anti-circumvention provisions rather than the CFAA route that failed in hiQ, arguing the defendants got around technical protection measures on two layers. If circumvention claims prove more durable than authorization claims, the practical line shifts from "was the data public" toward "what did you route around to reach it."
The case was still in progress as of early 2026, so treat any conclusion about it as provisional.
Public Doesn't Mean Unrestricted
Access is one question. What you may do with the data is several more:
- Copyright still applies. Publicly readable text is generally still protected. Extracting facts differs from republishing articles.
- Personal data carries its own rules. GDPR and similar regimes apply to personal data regardless of whether it was public, covering lawful basis, purpose limitation, and data-subject rights.
- Terms of service still bind those who agreed to them. This is the gap hiQ fell into. See Terms of Service violation.
- Circumvention is treated separately. Getting past a paywall, a login, or a technical protection measure is a different legal question than reading an open page, and increasingly the one plaintiffs lead with.
- Volume matters. Collection heavy enough to degrade a site invites claims that pure access questions wouldn't reach.
So, Is Public Data Scraping Legal?
The honest answer to is public data scraping legal is that collecting genuinely public data is, in most jurisdictions and on current US case law, the lowest-risk form of scraping, and still not risk-free. Risk rises sharply along a fairly predictable gradient: logging in, agreeing to terms, circumventing a protection measure, collecting personal data, republishing content, or scraping at a volume that harms the target.
Staying on the safe end mostly means staying logged out, respecting robots meta tag and robots.txt signals, rate-limiting yourself, collecting only fields you need, and getting advice before commercial-scale collection.
Related terms
Pagination (Web Scraping Pagination)
Learn what pagination means in web scraping, how numbered, next-link, and cursor patterns differ, and how to work through paged results without missing records.
Read more →XPath
Learn what XPath is, how it selects nodes by structure and content in HTML, and how it compares to CSS selectors for web scraping, with a syntax cheat sheet.
Read more →Web Scraping Pipeline
Learn what a web scraping pipeline is, the stages a production scraper runs through from fetch to storage, and why separating them keeps a scraper maintainable.
Read more →Web Unblocker
Extract data automatically, browse undetected, and beat anti-bot systems — all in one powerful tool.
Get started freeCommunity
Head over to our community where you can engage with us and our community directly.
Questions? Ask our team via live chat, join us on our official Slack community. We're always happy to help.
Join our Slack Community