Infinite Scroll Scraping
Infinite scroll scraping is extracting data from pages that load new content as you scroll down, rather than splitting results across numbered pages. Because each batch arrives through JavaScript after a scroll event, a scraper has to drive a real or headless browser, simulate scrolling, and wait for the new elements to render before reading them.
That last part is the whole difficulty. A plain HTTP request returns the initial HTML and nothing else, so it only ever sees the first batch of items no matter how many thousands exist below.
What Is Infinite Scroll?
Infinite scroll is a UI pattern where the page watches your scroll position and, as you approach the bottom, fires a background request for the next batch of results and appends them to the DOM. Social feeds, marketplace listings, image galleries, and search results all use it, because it keeps people scrolling instead of deciding whether to click "next".
Nothing about the URL changes while this happens. The page you started on simply grows.
Why a Plain HTTP Request Isn't Enough
A library like requests or a curl call fetches the HTML the server sent and stops. It doesn't run JavaScript, doesn't fire scroll events, and never triggers the fetch that loads batch two. You get the twenty items that shipped in the initial response and no indication that another two thousand were available.
Getting past that first batch requires JavaScript rendering in a real browser engine, which in practice means a headless browser driving the page programmatically.
Check for the Underlying API First
Before writing any scroll logic, open the browser's network tab and scroll once. The request that fires is usually a clean JSON endpoint with a cursor or offset parameter, something like /api/items?cursor=abc123&limit=20.
If you can call that endpoint directly, you skip the browser entirely: paginate by incrementing the cursor, get structured JSON instead of parsed HTML, and run an order of magnitude faster. Scroll simulation is the fallback for when the endpoint is signed, tied to a session, or otherwise not reusable.
How to Scrape an Infinite Scroll Page
When you do need the browser, the reliable loop is: count what's on the page, scroll, wait for the count to grow, repeat until it stops growing.
let previous = 0;
while (true) {
const count = await page.$$eval('.item', els => els.length);
if (count === previous) break; // nothing new arrived
previous = count;
await page.evaluate(() => window.scrollBy(0, document.body.scrollHeight));
await page
.waitForFunction(n => document.querySelectorAll('.item').length > n, {}, count)
.catch(() => {}); // end of feed, exit next pass
}
The detail that matters in any infinite scroll puppeteer script is the wait. A fixed sleep(2000) either wastes time on fast responses or truncates your dataset on slow ones. Waiting on an actual condition, such as the item count increasing, adapts to whatever the network does.
MrScraper handles this at the API level, where scroll automation is already built-in rather than something you script per target.
Infinite Scroll vs Pagination
| Infinite Scroll | Pagination | |
|---|---|---|
| Next batch trigger | Scroll position | A link or page number |
| URL changes | No | Yes, usually ?page=2 |
| Needs a browser | Yes, for scroll events and rendering | Often no, URLs can be fetched directly |
| Resumable | Hard, you replay scrolls from the top | Easy, jump straight to any page |
| Parallelizable | Poor, batches are sequential | Good, pages fetch independently |
| Knowing when to stop | Item count stops growing | Last page number or an empty result set |
On infinite scroll vs pagination, pagination is simply the easier target. If a site offers both, or still honors a ?page= parameter left over from an older layout, use it. See Pagination for handling that case.
Where It Breaks
- Virtualized lists. Some feeds recycle DOM nodes, removing items above the viewport as new ones load. The count never grows past a few dozen, and scrolling to the bottom leaves you with only the last batch. Extract after every scroll instead of once at the end.
- Lazy-loaded images. Image
srcattributes often populate only when a row enters the viewport, so a fast scroll leaves placeholders behind. - Duplicates and shifting order. Feeds reorder between requests, so the same item can appear twice. Deduplicate on a stable ID.
- Memory. A feed with tens of thousands of nodes will exhaust a headless browser tab. Cap the scroll count and checkpoint what you've collected.
- Behavioral checks. Instant jumps to the page bottom look nothing like human scrolling, which some anti-bot systems measure.
Related terms
Pagination (Web Scraping Pagination)
Learn what pagination means in web scraping, how numbered, next-link, and cursor patterns differ, and how to work through paged results without missing records.
Read more →XPath
Learn what XPath is, how it selects nodes by structure and content in HTML, and how it compares to CSS selectors for web scraping, with a syntax cheat sheet.
Read more →Public Data Scraping
Learn what public data scraping means, how hiQ v. LinkedIn shaped the legal picture, and why public access does not mean unrestricted use of the data.
Read more →Web Scraper API
Developer-friendly endpoints that return structured data in milliseconds.
Get started freeCommunity
Head over to our community where you can engage with us and our community directly.
Questions? Ask our team via live chat, join us on our official Slack community. We're always happy to help.
Join our Slack Community