Skip to content
← Back to glossary Scraping Basics

XPath

XPath is a query language for navigating and selecting nodes in an XML or HTML document by structure and content, not just by class names. That lets it do things CSS selectors can't, such as selecting an element purely by its visible text or walking up to a parent from a child match, at the cost of being more verbose and, on messy real-world pages, more fragile.

For scraping, it's the tool you reach for when the data you want has no useful class or ID, and the only reliable way to find it is by what it says or by what sits next to it.

What Is XPath?

XPath treats a document as a tree and gives you a path syntax for addressing any node in it. An expression describes a route through that tree, optionally filtered by conditions called predicates, and returns every node that matches.

Two features separate it from CSS selectors. XPath can match on text content, so //a[text()="Next"] finds a link by what it displays. And it can traverse in any direction through axes, so a match on a child can climb back up to its parent. CSS moves down and sideways only.

XPath Cheat Sheet

Expression Matches
//div Every div at any depth
/html/body/div A div at that exact absolute path
//div[@class="price"] A div whose class is exactly price
//div[contains(@class, "price")] A div whose class contains price
//a[text()="Next"] A link whose text is exactly "Next"
//a[contains(text(), "Next")] A link whose text contains "Next"
//span[@class="price"]/text() The text inside that span
//div[@id="main"]//a/@href The href values of links under #main
(//div[@class="item"])[1] The first matching item, counting from 1
//div[@class="item"][last()] The last matching item
//span[@class="price"]/parent::div The parent div of a matched span
//h2/following-sibling::p Paragraphs that follow an h2 at the same level
//div[@class="item"][.//span[@class="sale"]] Items that contain a sale badge somewhere inside

The last three rows are the ones with no CSS equivalent, and they're the reason XPath stays in the toolkit.

XPath vs CSS Selector

XPath CSS Selector
Select by text content Yes No
Traverse upward Yes, via parent:: and ancestor:: No
Sibling traversal Both directions Forward only
Readability Verbose, harder to scan Concise and familiar
Speed Slightly slower in most parsers Slightly faster
Browser support $x() in devtools, document.evaluate Native, querySelectorAll
Indexing Starts at 1 :nth-child also starts at 1

On xpath vs css selector, the working rule is to default to CSS and switch to XPath when you need text matching or upward traversal. Most scraping libraries accept both, so this is a per-selector decision rather than a project-wide one. MrScraper's PHP guide makes the same point from the other side: if XPath feels verbose, Symfony's DOMCrawler lets you extract content using familiar CSS selectors.

XPath in Web Scraping

In Python, lxml and Scrapy both take XPath directly:

python
from lxml import html
 
tree = html.fromstring(response.content)
prices = tree.xpath('//div[@class="product"]//span[@class="price"]/text()')
next_url = tree.xpath('//a[contains(text(), "Next")]/@href')

To test an expression before committing it to code, open devtools and run $x('//div[@class="price"]') in the console. It returns the matched nodes immediately, which is faster than a round trip through your scraper.

One warning about xpath web scraping workflows: don't use the browser's "Copy XPath" option. It produces absolute paths like /html/body/div[3]/div[2]/div[1]/span, which break the moment anything shifts anywhere above your target. Write the expression by hand against a stable attribute or text value instead.

Where XPath Breaks

  • Absolute paths. Positional paths encode the entire layout above your element. Any change breaks them.
  • Generated class names. Frameworks emit classes like css-1x2y3z that change on every build, so contains() on a partial class is a trap.
  • Phantom tbody. Browsers insert <tbody> into tables during parsing, so an XPath copied from devtools includes it while the raw HTML your parser sees does not.
  • Whitespace in text matches. text()="Next" fails on " Next ". Use normalize-space() instead.
  • Version limits. Most scraping libraries support XPath 1.0 only, so functions like matches() and ends-with() from 2.0 aren't available.
  • Shadow DOM. Content inside a shadow root is invisible to both XPath and CSS. When selectors of either kind become too brittle to maintain across many sites, selectorless scraping is the alternative that describes fields by meaning instead of by position.

Related terms

Web Unblocker

Extract data automatically, browse undetected, and beat anti-bot systems — all in one powerful tool.

Get started free

Community

Head over to our community where you can engage with us and our community directly.

Questions? Ask our team via live chat, join us on our official Slack community. We're always happy to help.

Join our Slack Community
Featured on CodeHype

Your choices

Cookie preferences

Necessary cookies keep your selection. Optional categories are disabled until you switch them on.

Strictly necessary

Remembers your privacy selection and keeps the site working.

Always on