Skip to content
← Back to glossary Scraping Basics

Infinite Scroll Scraping

Infinite scroll scraping is extracting data from pages that load new content as you scroll down, rather than splitting results across numbered pages. Because each batch arrives through JavaScript after a scroll event, a scraper has to drive a real or headless browser, simulate scrolling, and wait for the new elements to render before reading them.

That last part is the whole difficulty. A plain HTTP request returns the initial HTML and nothing else, so it only ever sees the first batch of items no matter how many thousands exist below.

What Is Infinite Scroll?

Infinite scroll is a UI pattern where the page watches your scroll position and, as you approach the bottom, fires a background request for the next batch of results and appends them to the DOM. Social feeds, marketplace listings, image galleries, and search results all use it, because it keeps people scrolling instead of deciding whether to click "next".

Nothing about the URL changes while this happens. The page you started on simply grows.

Why a Plain HTTP Request Isn't Enough

A library like requests or a curl call fetches the HTML the server sent and stops. It doesn't run JavaScript, doesn't fire scroll events, and never triggers the fetch that loads batch two. You get the twenty items that shipped in the initial response and no indication that another two thousand were available.

Getting past that first batch requires JavaScript rendering in a real browser engine, which in practice means a headless browser driving the page programmatically.

Check for the Underlying API First

Before writing any scroll logic, open the browser's network tab and scroll once. The request that fires is usually a clean JSON endpoint with a cursor or offset parameter, something like /api/items?cursor=abc123&limit=20.

If you can call that endpoint directly, you skip the browser entirely: paginate by incrementing the cursor, get structured JSON instead of parsed HTML, and run an order of magnitude faster. Scroll simulation is the fallback for when the endpoint is signed, tied to a session, or otherwise not reusable.

How to Scrape an Infinite Scroll Page

When you do need the browser, the reliable loop is: count what's on the page, scroll, wait for the count to grow, repeat until it stops growing.

js
let previous = 0;
 
while (true) {
  const count = await page.$$eval('.item', els => els.length);
  if (count === previous) break;           // nothing new arrived
  previous = count;
 
  await page.evaluate(() => window.scrollBy(0, document.body.scrollHeight));
 
  await page
    .waitForFunction(n => document.querySelectorAll('.item').length > n, {}, count)
    .catch(() => {});                      // end of feed, exit next pass
}

The detail that matters in any infinite scroll puppeteer script is the wait. A fixed sleep(2000) either wastes time on fast responses or truncates your dataset on slow ones. Waiting on an actual condition, such as the item count increasing, adapts to whatever the network does.

MrScraper handles this at the API level, where scroll automation is already built-in rather than something you script per target.

Infinite Scroll vs Pagination

Infinite Scroll Pagination
Next batch trigger Scroll position A link or page number
URL changes No Yes, usually ?page=2
Needs a browser Yes, for scroll events and rendering Often no, URLs can be fetched directly
Resumable Hard, you replay scrolls from the top Easy, jump straight to any page
Parallelizable Poor, batches are sequential Good, pages fetch independently
Knowing when to stop Item count stops growing Last page number or an empty result set

On infinite scroll vs pagination, pagination is simply the easier target. If a site offers both, or still honors a ?page= parameter left over from an older layout, use it. See Pagination for handling that case.

Where It Breaks

  • Virtualized lists. Some feeds recycle DOM nodes, removing items above the viewport as new ones load. The count never grows past a few dozen, and scrolling to the bottom leaves you with only the last batch. Extract after every scroll instead of once at the end.
  • Lazy-loaded images. Image src attributes often populate only when a row enters the viewport, so a fast scroll leaves placeholders behind.
  • Duplicates and shifting order. Feeds reorder between requests, so the same item can appear twice. Deduplicate on a stable ID.
  • Memory. A feed with tens of thousands of nodes will exhaust a headless browser tab. Cap the scroll count and checkpoint what you've collected.
  • Behavioral checks. Instant jumps to the page bottom look nothing like human scrolling, which some anti-bot systems measure.

Related terms

Web Scraper API

Developer-friendly endpoints that return structured data in milliseconds.

Get started free

Community

Head over to our community where you can engage with us and our community directly.

Questions? Ask our team via live chat, join us on our official Slack community. We're always happy to help.

Join our Slack Community
Featured on CodeHype

Your choices

Cookie preferences

Necessary cookies keep your selection. Optional categories are disabled until you switch them on.

Strictly necessary

Remembers your privacy selection and keeps the site working.

Always on