Skip to content
How to Scrape Dynamic Websites and Extract Tables in Python
Article

How to Scrape Dynamic Websites and Extract Tables in Python

Web Scraping

Learn how to scrape dynamic websites in Python using browser automation, pandas, BeautifulSoup, Requests, and APIs to extract structured table data.

By MrScraper Team 7 min read

To scrape dynamic websites in Python, use browser automation to render JavaScript tables. Then parse the HTML with BeautifulSoup or pandas. For static tables, pandas.read_html() or Requests and BeautifulSoup may be sufficient.

Why Table Scraping Matters

Unlike plain text on a page, HTML tables hold structured data in rows and columns. These map easily to spreadsheets or data frames. Whether you’re gathering:

  • Country statistics from Wikipedia
  • Financial market tables
  • Sports standings
  • Product feature matrices

…being able to extract tables programmatically saves you hours of manual effort. Python’s ecosystem gives you several tools that make this easier than you might expect.

Option 1: Quick and Easy with pandas.read_html()

One of the easiest ways to scrape tables in Python is with pandas’ built-in HTML table parser. [pandas.read_html()](https://pandas.pydata.org/docs/reference/api/pandas.read_html.html) reads all <table> elements from a URL or HTML string and returns them as DataFrames, ready to analyze.

Here’s how simple it can be:

import pandas as pd url = "https://en.wikipedia.org/wiki/List_of_countries_by_population_(United_Nations)" tables = pd.read_html(url) # Show how many tables were found print(f"Found {len(tables)} tables") # Work with the first table df = tables[0] print(df.head())
:--

Why this works:

  • Pandas uses lxml and BeautifulSoup under the hood to detect <table> structures and convert them into DataFrames.
  • You can pass a match parameter to filter only tables that contain specific text (e.g., a column header).

Pros: Minimal code, instant results.

Cons: Only works for static HTML; doesn’t handle JavaScript-rendered tables.

Option 2: BeautifulSoup + Requests: More Control

For scraping tables on pages where you need more control over parsing rows and cells, use Requests and BeautifulSoup

import requests from bs4 import BeautifulSoup import pandas as pd url = "https://example.com/table_page" response = requests.get(url) soup = BeautifulSoup(response.text, "html.parser") table = soup.find("table") # Find the first table # Extract header names headers = [th.get_text(strip=True) for th in table.find_all("th")] rows = [] for tr in table.find_all("tr"): cells = [td.get_text(strip=True) for td in tr.find_all("td")] if cells: rows.append(cells) df = pd.DataFrame(rows, columns=headers) print(df.head())
:--

What this does:

  1. Fetches the raw HTML using requests.
  2. Parses it with BeautifulSoup.
  3. Finds the <table> tag and extracts headers and cells.

Pros: Better error handling and control.

Cons: Requires more lines of code and understanding of HTML structure.

Option 3: Dynamic Tables with Browser Automation

To scrape dynamic websites, use browser automation when JavaScript creates the table after the initial page request. Standard HTTP requests may not capture that content, so Selenium can load the page in a real browser and expose the rendered HTML.

python
from bs4 import BeautifulSoup
import pandas as pd
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

def scrape_dynamic_table(url, table_selector="table", timeout=10):
    driver = webdriver.Chrome()
    try:
        driver.get(url)
        WebDriverWait(driver, timeout).until(
            EC.presence_of_element_located((By.CSS_SELECTOR, table_selector))
        )
        html = driver.page_source
        soup = BeautifulSoup(html, "html.parser")
        table = soup.select_one(table_selector)
        if table is None:
            raise ValueError(f"No table matched {table_selector!r}")
        df = pd.read_html(str(table))[0]
        return df
    finally:
        driver.quit()

df = scrape_dynamic_table("https://example.com/dynamic-table")
print(df.head())

This workflow opens a browser and waits for the table to load. It then parses the final HTML using BeautifulSoup and Pandas. The explicit wait is more reliable than a fixed delay and can be adjusted for slower pages.

  • Advantage: Selenium works with JavaScript-heavy sites whose tables appear only after rendering.
  • Trade-offs: Browser automation is slower. It needs a browser driver like ChromeDriver. It may need site-specific selectors. It may also need extra waiting logic.

Scrape Dynamic Websites with Scroll Cursors

This avoids duplicate records and gives the page time to load another batch. The row parsing approach in a practical table-scraping walkthrough also works once the browser has rendered the content.

python
from selenium import webdriver
from selenium.webdriver.common.by import By
import time

url = "https://example.com/dynamic-table"
driver = webdriver.Chrome()
rows_seen = set()

try:
    driver.get(url)
    last_height = 0
    while True:
        for row in driver.find_elements(By.CSS_SELECTOR, "table tbody tr"):
            cells = tuple(cell.text.strip() for cell in row.find_elements(By.TAG_NAME, "td"))
            if cells:
                rows_seen.add(cells)

        driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
        time.sleep(1.5)
        height = driver.execute_script("return document.body.scrollHeight")
        if height == last_height:
            break
        last_height = height
finally:
    driver.quit()

print(f"Collected {len(rows_seen)} unique rows")

That makes retries safe and prevents repeated records when a site overlaps adjacent pages.

Scrape Dynamic Websites with Playwright

It launches a browser, waits for the table to become available, and lets pandas parse the rendered markup. This pattern is useful when JavaScript creates the table after the initial response.

python
from io import StringIO
import pandas as pd
from playwright.sync_api import sync_playwright

url = "https://example.com/dynamic-table"

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    page.goto(url, wait_until="domcontentloaded")
    table = page.locator("table").first
    table.wait_for(state="visible")
    html = table.evaluate("element => element.outerHTML")
    browser.close()

df = pd.read_html(StringIO(html))[0]
print(df.head())

Option 4: Use APIs Behind Tables (When Available)

Before scraping HTML at all, it’s worth checking whether the table content is sourced from an API. Sites often load table data via XHR/Fetch requests. You can capture these API calls with your browser’s developer tools (Network tab). Then copy them using a simple Python request

import requests import pandas as pd api_url = "https://example.com/api/table-data" data = requests.get(api_url).json() df = pd.DataFrame(data["items"]) print(df.head())
:--

This method is often faster and cleaner than scraping HTML directly, and avoids HTML parsing complexities.

Tips for Reliable Table Scraping

  • Inspect the page’s HTML first. When you scrape dynamic websites, right-click the table and choose “Inspect Element” to understand its structure.
  • Choose the parser that fits the page. html.parser, lxml, and html5lib each involve trade-offs in speed and robustness.
  • Handle multiple tables deliberately. pd.read_html() returns a list if a page has more than one table. Choose the table you need by its index or by matching its content.
  • Respect robots.txt and the site’s Terms of Service. Check and comply with these policies before scraping large datasets.

MrScraper’s Table Extraction Support

For users who want to outsource this work to a managed service, MrScraper’s web scraping service can help. It offers robust, scalable table extraction capabilities:

  • Automated table detection and parsing: no need to write custom selectors.
  • JavaScript rendering support: handles sites where tables load dynamically.
  • Export options: get results in CSV or JSON format.
  • Proxy handling and anti-blocking logic: reduces the chances of request failures when scraping high-traffic sites.

Whether you’re scraping tables from e-commerce sites, public records, or research pages, MrScraper makes the process simple. It lets you focus on analyzing data instead of managing scraping infrastructure.

Conclusion

Web scraping tables in Python is easier than many people realize, thanks to a rich set of libraries like pandas, BeautifulSoup, requests, and Selenium. For simple static tables, Pandas’ read_html() can pull data into a DataFrame with just a couple of lines of code. For more complex scenarios, BeautifulSoup and browser automation give you precision and flexibility.

With these techniques, you can extract structured table data from many websites. You can turn HTML into usable datasets for analytics, reporting, or machine learning. You can also do it with Python code.

What We Learned

This approach turns the guide’s individual techniques into a repeatable workflow. Start by checking the delivered HTML. Then inspect network requests when the table loads later. Use browser automation only if neither source is enough.

  • Choose the simplest source that contains the complete table data.
  • Normalize headers and values before combining results from multiple pages.
  • Validate row counts, required columns, and data types after each run.
  • Save the source URL, retrieval time, and extraction method with the dataset.
  • Check robots.txt and the site’s terms before collecting data at scale.

The result is more than a scraped table. It is a traceable dataset with a clear extraction path. It includes checks to catch silent failures. The workflow can adapt when a website changes.

Continue with a Table-Scraping Quickstart

Explore a practical starting point for managed table extraction when Python code is not a good fit.

Get Started

Summarize this post

Open it in your assistant of choice with the prompt ready to send.

Take a Taste of Easy Scraping!

Your choices

Cookie preferences

Necessary cookies keep your selection. Optional categories are disabled until you switch them on.

Strictly necessary

Remembers your privacy selection and keeps the site working.

Always on