How to Scrape Dynamic Websites and Extract Tables in Python
Web ScrapingLearn how to scrape dynamic websites in Python using browser automation, pandas, BeautifulSoup, Requests, and APIs to extract structured table data.
To scrape dynamic websites in Python, use browser automation to render JavaScript tables. Then parse the HTML with BeautifulSoup or pandas. For static tables, pandas.read_html() or Requests and BeautifulSoup may be sufficient.
Why Table Scraping Matters
Unlike plain text on a page, HTML tables hold structured data in rows and columns. These map easily to spreadsheets or data frames. Whether you’re gathering:
- Country statistics from Wikipedia
- Financial market tables
- Sports standings
- Product feature matrices
…being able to extract tables programmatically saves you hours of manual effort. Python’s ecosystem gives you several tools that make this easier than you might expect.
Option 1: Quick and Easy with pandas.read_html()
One of the easiest ways to scrape tables in Python is with pandas’ built-in HTML table parser. [pandas.read_html()](https://pandas.pydata.org/docs/reference/api/pandas.read_html.html) reads all <table> elements from a URL or HTML string and returns them as DataFrames, ready to analyze.
Here’s how simple it can be:
import pandas as pd url = "https://en.wikipedia.org/wiki/List_of_countries_by_population_(United_Nations)" tables = pd.read_html(url) # Show how many tables were found print(f"Found {len(tables)} tables") # Work with the first table df = tables[0] print(df.head()) |
|---|
| :-- |
Why this works:
- Pandas uses
lxmlandBeautifulSoupunder the hood to detect<table>structures and convert them into DataFrames. - You can pass a
matchparameter to filter only tables that contain specific text (e.g., a column header).
Pros: Minimal code, instant results.
Cons: Only works for static HTML; doesn’t handle JavaScript-rendered tables.
Option 2: BeautifulSoup + Requests: More Control
For scraping tables on pages where you need more control over parsing rows and cells, use Requests and BeautifulSoup
import requests from bs4 import BeautifulSoup import pandas as pd url = "https://example.com/table_page" response = requests.get(url) soup = BeautifulSoup(response.text, "html.parser") table = soup.find("table") # Find the first table # Extract header names headers = [th.get_text(strip=True) for th in table.find_all("th")] rows = [] for tr in table.find_all("tr"): cells = [td.get_text(strip=True) for td in tr.find_all("td")] if cells: rows.append(cells) df = pd.DataFrame(rows, columns=headers) print(df.head()) |
|---|
| :-- |
What this does:
- Fetches the raw HTML using
requests. - Parses it with
BeautifulSoup. - Finds the
<table>tag and extracts headers and cells.
Pros: Better error handling and control.
Cons: Requires more lines of code and understanding of HTML structure.
Option 3: Dynamic Tables with Browser Automation
To scrape dynamic websites, use browser automation when JavaScript creates the table after the initial page request. Standard HTTP requests may not capture that content, so Selenium can load the page in a real browser and expose the rendered HTML.
from bs4 import BeautifulSoup
import pandas as pd
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
def scrape_dynamic_table(url, table_selector="table", timeout=10):
driver = webdriver.Chrome()
try:
driver.get(url)
WebDriverWait(driver, timeout).until(
EC.presence_of_element_located((By.CSS_SELECTOR, table_selector))
)
html = driver.page_source
soup = BeautifulSoup(html, "html.parser")
table = soup.select_one(table_selector)
if table is None:
raise ValueError(f"No table matched {table_selector!r}")
df = pd.read_html(str(table))[0]
return df
finally:
driver.quit()
df = scrape_dynamic_table("https://example.com/dynamic-table")
print(df.head())
This workflow opens a browser and waits for the table to load. It then parses the final HTML using BeautifulSoup and Pandas. The explicit wait is more reliable than a fixed delay and can be adjusted for slower pages.
- Advantage: Selenium works with JavaScript-heavy sites whose tables appear only after rendering.
- Trade-offs: Browser automation is slower. It needs a browser driver like ChromeDriver. It may need site-specific selectors. It may also need extra waiting logic.
Scrape Dynamic Websites with Scroll Cursors
This avoids duplicate records and gives the page time to load another batch. The row parsing approach in a practical table-scraping walkthrough also works once the browser has rendered the content.
from selenium import webdriver
from selenium.webdriver.common.by import By
import time
url = "https://example.com/dynamic-table"
driver = webdriver.Chrome()
rows_seen = set()
try:
driver.get(url)
last_height = 0
while True:
for row in driver.find_elements(By.CSS_SELECTOR, "table tbody tr"):
cells = tuple(cell.text.strip() for cell in row.find_elements(By.TAG_NAME, "td"))
if cells:
rows_seen.add(cells)
driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
time.sleep(1.5)
height = driver.execute_script("return document.body.scrollHeight")
if height == last_height:
break
last_height = height
finally:
driver.quit()
print(f"Collected {len(rows_seen)} unique rows")
That makes retries safe and prevents repeated records when a site overlaps adjacent pages.
Scrape Dynamic Websites with Playwright
It launches a browser, waits for the table to become available, and lets pandas parse the rendered markup. This pattern is useful when JavaScript creates the table after the initial response.
from io import StringIO
import pandas as pd
from playwright.sync_api import sync_playwright
url = "https://example.com/dynamic-table"
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
page.goto(url, wait_until="domcontentloaded")
table = page.locator("table").first
table.wait_for(state="visible")
html = table.evaluate("element => element.outerHTML")
browser.close()
df = pd.read_html(StringIO(html))[0]
print(df.head())
Option 4: Use APIs Behind Tables (When Available)
Before scraping HTML at all, it’s worth checking whether the table content is sourced from an API. Sites often load table data via XHR/Fetch requests. You can capture these API calls with your browser’s developer tools (Network tab). Then copy them using a simple Python request
import requests import pandas as pd api_url = "https://example.com/api/table-data" data = requests.get(api_url).json() df = pd.DataFrame(data["items"]) print(df.head()) |
|---|
| :-- |
This method is often faster and cleaner than scraping HTML directly, and avoids HTML parsing complexities.
Tips for Reliable Table Scraping
- Inspect the page’s HTML first. When you scrape dynamic websites, right-click the table and choose “Inspect Element” to understand its structure.
- Choose the parser that fits the page. html.parser, lxml, and html5lib each involve trade-offs in speed and robustness.
- Handle multiple tables deliberately. pd.read_html() returns a list if a page has more than one table. Choose the table you need by its index or by matching its content.
- Respect robots.txt and the site’s Terms of Service. Check and comply with these policies before scraping large datasets.
MrScraper’s Table Extraction Support
For users who want to outsource this work to a managed service, MrScraper’s web scraping service can help. It offers robust, scalable table extraction capabilities:
- Automated table detection and parsing: no need to write custom selectors.
- JavaScript rendering support: handles sites where tables load dynamically.
- Export options: get results in CSV or JSON format.
- Proxy handling and anti-blocking logic: reduces the chances of request failures when scraping high-traffic sites.
Whether you’re scraping tables from e-commerce sites, public records, or research pages, MrScraper makes the process simple. It lets you focus on analyzing data instead of managing scraping infrastructure.
Conclusion
Web scraping tables in Python is easier than many people realize, thanks to a rich set of libraries like pandas, BeautifulSoup, requests, and Selenium. For simple static tables, Pandas’ read_html() can pull data into a DataFrame with just a couple of lines of code. For more complex scenarios, BeautifulSoup and browser automation give you precision and flexibility.
With these techniques, you can extract structured table data from many websites. You can turn HTML into usable datasets for analytics, reporting, or machine learning. You can also do it with Python code.
What We Learned
This approach turns the guide’s individual techniques into a repeatable workflow. Start by checking the delivered HTML. Then inspect network requests when the table loads later. Use browser automation only if neither source is enough.
- Choose the simplest source that contains the complete table data.
- Normalize headers and values before combining results from multiple pages.
- Validate row counts, required columns, and data types after each run.
- Save the source URL, retrieval time, and extraction method with the dataset.
- Check robots.txt and the site’s terms before collecting data at scale.
The result is more than a scraped table. It is a traceable dataset with a clear extraction path. It includes checks to catch silent failures. The workflow can adapt when a website changes.
Continue with a Table-Scraping Quickstart
Explore a practical starting point for managed table extraction when Python code is not a good fit.
Summarize this post
Open it in your assistant of choice with the prompt ready to send.
Take a Taste of Easy Scraping!
Find more insights here
The Ultimate Web Crawlers List: 15 Tools for Every Data Need
Compare the best web crawlers for 2026. Learn the difference between open-source, managed APIs, and…

Web Scraping MCP Server: Giving AI Agents Direct Access to Live Web Data
Learn how a Web Scraping MCP Server gives AI agents live web access while reducing token costs by 87…

Scaling E-commerce Competitive Intelligence with Automated Data Harvesting
Scale e-commerce data harvesting with residential proxies and AI. Learn how modern data extraction s…