How to Scrape Dynamic Websites: 4 Ways to Find All URLs on a Domain
Web ScrapingLearn how to scrape dynamic websites and find all URLs on a domain with MrScraper, XML sitemaps, Google Search Console, or a Python crawler.
To scrape dynamic sites and find all domain URLs, use MrScraper, an XML sitemap, Google Search Console, or a Python crawler. Choose the method that matches your access, technical skills, and need for control.
Why Finding All URLs is Important
To scrape dynamic websites, first understand why discovering every URL on a domain matters. A web scraping tool can help extract URLs from a website, including complex sites. This guide presents several ways to find them, with particular attention to using a web scraping tool.
In today's digital age, understanding a website's structure is valuable for web developers, SEO experts, and digital marketers. Whether you are conducting a site audit, analyzing competitors, or preparing for a site migration, identifying all URLs can reveal important information about the site and its content.
- SEO audits can uncover hidden and orphan pages and help verify that URLs are properly indexed.
- A content inventory creates a complete list of content assets for repurposing, updating, or migration.
- Competitor analysis can reveal how another site is structured and provide insight into its content strategy.
- A broken-link check can identify links that need to be fixed and may be harming SEO.
Method 1: Using MrScraper to Find All URLs on a Domain
To scrape dynamic websites and collect their URLs, use the platform's guided project workflow. It is designed to extract URLs from a domain using a no-code interface. It works on complex websites. It does not require programming knowledge.
- Sign up for an account. Create an account before starting if you do not already have one. The interface provides access to the scraping features needed for this workflow.
- Create a new project. Enter the domain you want to inspect, then select the option to scrape all URLs.
- Configure the scraping settings. Adjust options such as the crawl depth and URL filters to match the scope of your project. These settings help limit the crawl to the pages and paths you want to review.
- Run the scraper. Start the project after confirming its settings. The scraper crawls the domain and produces a list of the URLs it finds.
- Export the results. When the crawl is complete, export the URL list as CSV or JSON. Use it for analysis, filtering, or another workflow.
This method is useful when you want a guided alternative to writing a crawler yourself. Its AI-powered, no-code features make the workflow easy for non-technical users. Its integrations can automate recurring scraping tasks and reduce manual work. Review the selected domain, crawl depth, and URL filters before you run the project. This helps the exported results match your intended scope.
Scrape Dynamic Websites
import { chromium } from "playwright";
const browser = await chromium.launch();
const page = await browser.newPage();
await page.goto("https://example.com", { waitUntil: "networkidle" });
await page.evaluate(() => window.scrollTo(0, document.body.scrollHeight));
const links = await page.locator("a[href]").evaluateAll((anchors) =>
[...new Set(anchors.map((a) => new URL(a.href, location.href).href))]
);
console.log(links);
await browser.close();
Google Search Central’s JavaScript SEO basics is a useful reference for understanding how crawlers process rendered content.
Method 2: Using XML Sitemaps
To scrape dynamic websites, start by checking whether the domain provides an XML sitemap. This is a common way to find the URLs the site has indexed. Most sitemaps are available at the domain root under the standard sitemap.xml path. First, open the domain’s XML sitemap. Then extract its URL entries. You can copy the URLs manually, or use Python to download and parse the XML document. The extracted list gives you the sitemap’s indexed URLs for further review or crawling. For programmatic extraction, see the previous guide, Parsing XML with Python: A Comprehensive Guide. It explains how to read XML files with Python and apply the same approach to sitemap data.
Method 3: Google Search Console
Answer: To scrape dynamic websites, use Google Search Console to retrieve the URLs Google has indexed for the domain. If you can access the domain’s property, the service provides an exportable list. It can help you identify indexed pages.
- Sign in to Google Search Console and select the property corresponding to your domain.
- Open the Coverage report under the Index section. Review the URLs that Google has indexed.
- Export the URL data from Google Search Console to obtain a comprehensive list of indexed URLs.
Method 4: Manual Crawling with Python
To scrape dynamic websites, first check whether the links appear in the HTML returned by the server. The crawler below provides a hands-on Python method with direct control over which pages are visited. It follows links on the same host. It converts relative links into absolute URLs. It records each discovered URL only once. If a page adds links only after JavaScript runs, requests will not see those links without a browser-rendering step. The script starts at the supplied domain, continues through queued pages, and prints the URLs it finds. Request failures and unsuccessful HTTP responses are reported without stopping the rest of the crawl. Requests fetches each page, BeautifulSoup parses its HTML, and urljoin resolves relative links. You can save the returned set to a file or pass it to later processing.
import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin, urlparse
def find_urls(domain):
start_url = domain.rstrip("/")
allowed_host = urlparse(start_url).netloc.lower()
urls = {start_url}
to_crawl = [start_url]
while to_crawl:
url = to_crawl.pop(0)
try:
response = requests.get(url, timeout=10)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
for link in soup.find_all("a", href=True):
full_url = urljoin(url, link["href"])
parsed_url = urlparse(full_url)
if (parsed_url.scheme in ("http", "https")
and parsed_url.netloc.lower() == allowed_host
and full_url not in urls):
urls.add(full_url)
to_crawl.append(full_url)
except requests.exceptions.RequestException as error:
print(f"Failed to crawl {url}: {error}")
return urls
domain = "https://www.example.com"
urls = find_urls(domain)
for url in sorted(urls):
print(url)
The crawl is limited to the starting host, so links to other websites are not added. Review or filter the collected URLs before using them in another process.
Scrape Dynamic Websites with Playwright
Sitemap and HTML crawling can be complemented when a page adds URLs after load. Keep the crawl bounded by the target hostname and a page limit.
import asyncio
from urllib.parse import urljoin, urlparse
from playwright.async_api import async_playwright
async def extract_urls(start_url, limit=20):
host = urlparse(start_url).netloc
seen, queue, found = {start_url}, [start_url], set()
async with async_playwright() as p:
browser = await p.chromium.launch()
page = await browser.new_page()
while queue and len(seen) <= limit:
current = queue.pop(0)
await page.goto(current, wait_until="networkidle")
for href in await page.locator("a[href]").evaluate_all(
"els => els.map(e => e.href)"
):
url = urljoin(current, href).split("#", 1)[0]
if urlparse(url).netloc == host:
found.add(url)
if url not in seen:
seen.add(url)
queue.append(url)
await browser.close()
return sorted(found)
print(asyncio.run(extract_urls("https://example.com")))
The Playwright documentation is a useful reference for navigation waits, locator evaluation, and browser lifecycle management.
Conclusion
Finding all URLs on a domain is an important task for web development, SEO, and digital marketing. To scrape dynamic websites, use a web scraping tool, XML sitemaps, and Google Search Console. These methods work together to help you find URLs. Each approach offers different benefits. If you prefer hands-on control, building a crawler in Python lets you manage the crawling process yourself. Choose the method that best fits your workflow and the site's scope.
What We Learned
Use the coverage triangulation pattern. Start with a crawl. Compare the crawl URL list with the sitemap. Then check indexed URLs in Search Console. Investigate differences instead of assuming one source is complete. For a programmable workflow, a Python crawler can provide a repeatable baseline. Other sources can reveal URLs the crawl may not find. This approach also gives the final inventory a clear purpose. The next step may be an audit, migration, content review, or broken-link check.
- Choose a discovery method that matches the site and the information you need.
- Compare independent URL sources before treating the inventory as complete.
- Export the results so the inventory can support later analysis.
Start Finding URLs with MrScraper
Use the quickstart resources to plan a URL discovery workflow with MrScraper. Review the steps for extracting URLs from a domain.
Summarize this post
Open it in your assistant of choice with the prompt ready to send.
Take a Taste of Easy Scraping!
Find more insights here

Web Scraping MCP Server: Giving AI Agents Direct Access to Live Web Data
Learn how a Web Scraping MCP Server gives AI agents live web access while reducing token costs by 87…

Scaling E-commerce Competitive Intelligence with Automated Data Harvesting
Scale e-commerce data harvesting with residential proxies and AI. Learn how modern data extraction s…

Scaling Data Extraction via AI-Driven Dynamic Selectors
Learn how AI-driven dynamic selectors and residential proxies reduce web scraping maintenance costs…