Robots Meta Tag (noindex/nofollow)
A robots meta tag is an HTML <meta> element that tells crawlers whether to index a page and whether to follow its links, with noindex and nofollow being the two most common directives. Unlike robots.txt, which blocks crawling at the site or path level before a page is ever fetched, the meta tag only takes effect once the page has already been retrieved and parsed.
That ordering is the point people miss. A crawler must download a page to read its robots meta tag, so the tag controls what happens to the content afterward, not whether the request is made.
What Is a Robots Meta Tag?
It sits in the <head> and applies to that page alone:
<meta name="robots" content="noindex, nofollow">
The name attribute targets which crawlers obey it. robots addresses all of them; a specific token addresses one, so <meta name="googlebot" content="noindex"> applies to Google while leaving others unaffected. The content attribute holds a comma-separated list of directives.
noindex, nofollow, and the Other Directives
noindex: don't include this page in search results. The page can still be crawled, and its links still followed.nofollow: don't follow links on this page or pass ranking signals through them.none: shorthand fornoindex, nofollow.noarchive: don't show a cached copy.nosnippet: don't show a text snippet or video preview in results.max-snippet:[n],max-image-preview:[setting],max-video-preview:[n]: cap how much of the page appears in results.noimageindex: don't index images on the page.unavailable_after:[date]: stop showing the page after a given date. Common meta robots noindex example cases are internal search results, thin tag archives, staging pages, checkout and account screens, and printer-friendly duplicates. The default when no tag is present isindex, follow, so there's no reason to state that explicitly.
One frequent mistake is worth naming: blocking a page in robots.txt and adding noindex prevents deindexing rather than ensuring it. If the crawler is not allowed to fetch the page, it never sees the noindex, and the URL can persist in results on the strength of inbound links alone.
Robots Meta Tag vs robots.txt
| Robots Meta Tag | robots.txt | |
|---|---|---|
| Location | In each page's <head> |
One file at the domain root |
| Scope | That single page | Paths, directories, or the whole site |
| Takes effect | After the page is fetched and parsed | Before the request is made |
| Controls | Indexing, link following, snippet display | Crawling and request volume |
| Saves bandwidth | No, the page is downloaded either way | Yes, the fetch never happens |
| Right tool for | Keeping a fetched page out of search results | Keeping crawlers away from sections entirely |
On robots meta tag vs robots.txt, the clean rule is: robots.txt controls access, the meta tag controls treatment. If you want a page crawled but not listed, use the meta tag. If you want a whole directory left alone, use robots.txt.
There's also the X-Robots-Tag HTTP header, which carries the same directives in the response rather than the HTML. That's how you apply noindex to non-HTML files like PDFs and images, where there's no <head> to put a tag in.
What It Means for Scrapers
Robots directives are instructions to indexers, not access controls. They don't authenticate, rate-limit, or block anything, and a scraper that ignores them is not bypassing a technical protection measure the way credential stuffing or paywall circumvention would be.
That said, they're a clear statement of intent from the site owner, and the SEO-oriented directives are mostly irrelevant to scraping anyway. Whether a page appears in Google has no bearing on whether you can extract data from it. The signals that actually matter for a web crawler are robots.txt disallow rules, crawl-delay, and the site's terms, alongside your own crawl depth and rate limits.
With the emergence of llms.txt and AI-specific crawler tokens, site owners now have three separate layers for expressing crawler preferences: robots.txt for access, the robots meta tag for indexing treatment, and llms.txt for what AI systems should read. None of them enforce anything on their own.
Related terms
Terms of Service (ToS) Violation
Learn what a Terms of Service violation means for web scraping, why it is a contract question rather than a criminal one, and how it differs from copyright claims.
Read more →Behavioral Biometrics
Behavioral biometrics analyzes mouse movements and keystroke dynamics to detect bots. Learn how anti-bot systems use behavioral signals to flag scrapers.
Read more →Good Bot vs Bad Bot
Learn the difference between good bot vs bad bot traffic, how modern firewalls classify automated crawlers, and how scrapers navigate anti-bot detection.
Read more →Web Unblocker
Extract data automatically, browse undetected, and beat anti-bot systems — all in one powerful tool.
Get started freeCommunity
Head over to our community where you can engage with us and our community directly.
Questions? Ask our team via live chat, join us on our official Slack community. We're always happy to help.
Join our Slack Community