Skip to content
← Back to glossary Scraping Basics

Pagination (Web Scraping Pagination)

Pagination in web scraping means systematically working through a site's or API's paged results by incrementing a URL parameter, following a "next" link, or advancing a cursor token, rather than stopping after the first page. Which approach you need depends entirely on the pattern the target uses.

Getting this wrong is the most common reason a scraper quietly returns 20 records when 4,000 were available. The first page usually works on the first try, which is exactly what makes the problem easy to miss.

What Is Pagination?

Pagination is the practice of splitting a large result set into smaller chunks that load separately. What is pagination on the web usually means numbered page links at the bottom of a listing, but the same idea covers "load more" buttons, API cursors, and feeds that extend as you scroll.

For a scraper, the useful question isn't what it looks like but how the next chunk is requested. That determines whether you can fetch pages directly or have to drive a browser.

The Four Patterns

Numbered or offset URLs. The page number lives in the URL: ?page=2, ?p=2, ?offset=40&limit=20, or a path segment like /products/page/2. This is the easiest case by far, because every page has its own address, requests can run in parallel, and a failed page is retried on its own.

Next-link following. No predictable parameter, just a "next" anchor whose href you extract and follow, one page at a time. Sequential by nature, since you can't know page five's URL until you've fetched page four.

Load-more buttons. Content extends in place when a button is clicked. The click fires a background request, so the practical move is to find that request in the network tab and call it directly rather than automating clicks.

Cursor or token based. Each response carries an opaque token pointing at the next batch. Common in APIs and modern feeds.

API Pagination: Offset vs Cursor

Offset / Page Cursor / Token
Next request built from A number you increment A token returned in the previous response
Jump to an arbitrary page Yes No, strictly sequential
Parallel fetching Yes No
Behavior on live data Items shift between pages, causing duplicates and gaps Stable, the cursor marks a fixed position
Performance at depth Degrades, the database skips N rows Consistent
Resumability Store the page number Store the cursor

Cursor pagination is now the default in most large APIs because of that fourth row. With offset pagination on a feed where new items arrive constantly, everything shifts down by one between requests, so an item that was last on page one becomes first on page two and you collect it twice, while something else is skipped entirely. A cursor points at a record rather than a position, so insertions don't disturb it.

The trade-off is that api pagination by cursor can't be parallelized. Offset pagination lets you fire pages 1 through 50 at once; cursors force one request at a time. On static catalogs, offset is often the faster choice for exactly this reason.

Pagination vs Infinite Scroll

These get conflated, and the difference is mechanical. With pagination, the next batch is requested by a link, a button, or a parameter you control, and each batch usually has its own URL. With infinite scroll, the next batch is triggered by scroll position, the URL never changes, and a plain HTTP request only ever sees the first batch.

The practical consequence is tooling. Paginated pages can typically be fetched with a simple HTTP client, no browser required. Infinite scroll needs a headless browser to generate the scroll events, unless you can call the underlying endpoint directly. Where a site offers both, or where an older ?page= parameter still works alongside a scrolling layout, take the paginated route.

Knowing When to Stop

To scrape paginated pages reliably, you need a termination condition that isn't just a page count:

  • The next link is absent or disabled.
  • The response returns an empty result set.
  • The page returns a 404 or redirects to page one.
  • Items repeat content you already collected.
  • You reach a last-page number parsed from the pager itself. Check more than one of these. Sites frequently return a valid-looking page with zero results rather than a 404, and some return page one again for any page number past the end, which turns a naive loop into an infinite one.

Common Pitfalls

  • Hard caps. Many sites stop serving results past page 100 or so regardless of the stated total. The fix is narrowing your query with filters or sort order, not paging harder.
  • Shifting results. On live listings, sort by something stable such as ID or creation date rather than relevance or popularity.
  • Off-by-one parameters. Some APIs count pages from 0, others from 1. Verify against a known record instead of assuming.
  • Silent truncation. A limit parameter set above the server's maximum is often ignored rather than rejected, capping your batch at the default size.
  • Parallel fetching without limits. Numbered URLs make concurrency tempting. Fifty simultaneous requests is also a reliable way to get rate-limited.
  • No checkpointing. A run that dies on page 380 of 400 should resume, not restart. Persist the last completed page or cursor.

Related terms

Web Unblocker

Extract data automatically, browse undetected, and beat anti-bot systems — all in one powerful tool.

Get started free

Community

Head over to our community where you can engage with us and our community directly.

Questions? Ask our team via live chat, join us on our official Slack community. We're always happy to help.

Join our Slack Community
Featured on CodeHype

Your choices

Cookie preferences

Necessary cookies keep your selection. Optional categories are disabled until you switch them on.

Strictly necessary

Remembers your privacy selection and keeps the site working.

Always on