XPath
XPath is a query language for navigating and selecting nodes in an XML or HTML document by structure and content, not just by class names. That lets it do things CSS selectors can't, such as selecting an element purely by its visible text or walking up to a parent from a child match, at the cost of being more verbose and, on messy real-world pages, more fragile.
For scraping, it's the tool you reach for when the data you want has no useful class or ID, and the only reliable way to find it is by what it says or by what sits next to it.
What Is XPath?
XPath treats a document as a tree and gives you a path syntax for addressing any node in it. An expression describes a route through that tree, optionally filtered by conditions called predicates, and returns every node that matches.
Two features separate it from CSS selectors. XPath can match on text content, so //a[text()="Next"] finds a link by what it displays. And it can traverse in any direction through axes, so a match on a child can climb back up to its parent. CSS moves down and sideways only.
XPath Cheat Sheet
| Expression | Matches |
|---|---|
//div |
Every div at any depth |
/html/body/div |
A div at that exact absolute path |
//div[@class="price"] |
A div whose class is exactly price |
//div[contains(@class, "price")] |
A div whose class contains price |
//a[text()="Next"] |
A link whose text is exactly "Next" |
//a[contains(text(), "Next")] |
A link whose text contains "Next" |
//span[@class="price"]/text() |
The text inside that span |
//div[@id="main"]//a/@href |
The href values of links under #main |
(//div[@class="item"])[1] |
The first matching item, counting from 1 |
//div[@class="item"][last()] |
The last matching item |
//span[@class="price"]/parent::div |
The parent div of a matched span |
//h2/following-sibling::p |
Paragraphs that follow an h2 at the same level |
//div[@class="item"][.//span[@class="sale"]] |
Items that contain a sale badge somewhere inside |
The last three rows are the ones with no CSS equivalent, and they're the reason XPath stays in the toolkit.
XPath vs CSS Selector
| XPath | CSS Selector | |
|---|---|---|
| Select by text content | Yes | No |
| Traverse upward | Yes, via parent:: and ancestor:: |
No |
| Sibling traversal | Both directions | Forward only |
| Readability | Verbose, harder to scan | Concise and familiar |
| Speed | Slightly slower in most parsers | Slightly faster |
| Browser support | $x() in devtools, document.evaluate |
Native, querySelectorAll |
| Indexing | Starts at 1 | :nth-child also starts at 1 |
On xpath vs css selector, the working rule is to default to CSS and switch to XPath when you need text matching or upward traversal. Most scraping libraries accept both, so this is a per-selector decision rather than a project-wide one. MrScraper's PHP guide makes the same point from the other side: if XPath feels verbose, Symfony's DOMCrawler lets you extract content using familiar CSS selectors.
XPath in Web Scraping
In Python, lxml and Scrapy both take XPath directly:
from lxml import html
tree = html.fromstring(response.content)
prices = tree.xpath('//div[@class="product"]//span[@class="price"]/text()')
next_url = tree.xpath('//a[contains(text(), "Next")]/@href')
To test an expression before committing it to code, open devtools and run $x('//div[@class="price"]') in the console. It returns the matched nodes immediately, which is faster than a round trip through your scraper.
One warning about xpath web scraping workflows: don't use the browser's "Copy XPath" option. It produces absolute paths like /html/body/div[3]/div[2]/div[1]/span, which break the moment anything shifts anywhere above your target. Write the expression by hand against a stable attribute or text value instead.
Where XPath Breaks
- Absolute paths. Positional paths encode the entire layout above your element. Any change breaks them.
- Generated class names. Frameworks emit classes like
css-1x2y3zthat change on every build, socontains()on a partial class is a trap. - Phantom
tbody. Browsers insert<tbody>into tables during parsing, so an XPath copied from devtools includes it while the raw HTML your parser sees does not. - Whitespace in text matches.
text()="Next"fails on" Next ". Usenormalize-space()instead. - Version limits. Most scraping libraries support XPath 1.0 only, so functions like
matches()andends-with()from 2.0 aren't available. - Shadow DOM. Content inside a shadow root is invisible to both XPath and CSS. When selectors of either kind become too brittle to maintain across many sites, selectorless scraping is the alternative that describes fields by meaning instead of by position.
Related terms
Pagination (Web Scraping Pagination)
Learn what pagination means in web scraping, how numbered, next-link, and cursor patterns differ, and how to work through paged results without missing records.
Read more →Public Data Scraping
Learn what public data scraping means, how hiQ v. LinkedIn shaped the legal picture, and why public access does not mean unrestricted use of the data.
Read more →Web Scraping Pipeline
Learn what a web scraping pipeline is, the stages a production scraper runs through from fetch to storage, and why separating them keeps a scraper maintainable.
Read more →Web Unblocker
Extract data automatically, browse undetected, and beat anti-bot systems — all in one powerful tool.
Get started freeCommunity
Head over to our community where you can engage with us and our community directly.
Questions? Ask our team via live chat, join us on our official Slack community. We're always happy to help.
Join our Slack Community