Skip to content
Free tool · No signup · Nothing to install

Website to Text: Convert Any Page to Clean Markdown

Copying from a browser brings the navigation and the cookie banner with it. Enter a page address instead and get clean Markdown or plain text back: the structure kept, the boilerplate stripped, ready for an AI model, a RAG pipeline or a reader.

Keeps headings, lists, links, tables and code. Drops menus, scripts, styles and footers.

How to convert a website to text

  1. 1

    Enter a page address

    Any public page works: an article, a docs page, a product page or a changelog.

  2. 2

    Press Convert to text

    The page is fetched, its main content is found, and the markup is rewritten as Markdown.

  3. 3

    Copy or download

    Switch between Markdown and plain text, then copy it or save it as a .md or .txt file.

What is kept and what is removed

  • Headings

    h1 to h6 become # to ######, so the outline of the page survives.

  • Links and images

    Links keep their absolute URLs. Images with alt text are kept, and inline data images are dropped.

  • Lists, tables and quotes

    Nested lists keep their indent, tables become pipe tables and blockquotes keep their marker.

  • Code

    Code blocks become fenced blocks with their whitespace intact, and inline code keeps its backticks.

  • Removed: page chrome

    Navigation, sidebars, footers, forms and buttons are dropped. When the page has a main or article element, only that is read.

  • Removed: code that is not content

    Scripts, styles, SVG, iframes and HTML comments never reach the output.

The same page as HTML and as Markdown

Stripping scripts and styling removes most of the byte weight of a page without losing meaning. One real page, measured before and after conversion:

Raw HTML

236 KB

about 58,868 tokens

Markdown

49 KB

about 12,195 tokens

Smaller by

79%

same information

Measured on 28 September 2026 with this converter on the English Wikipedia article on web scraping. Tokens are estimated at four characters each.

Copy-paste versus this converter versus a scraping API.
Aspect Browser copy-paste This free converter A scraping API
Heading structure Lost when pasted as plain text Kept as # to ###### Kept, or returned as structured fields
Navigation, footers, scripts Come along with the selection Dropped before conversion Dropped before conversion
JavaScript-rendered content Included, since the browser ran it Missing: only the server HTML is read Included when it renders in a headless browser
Many pages or a schedule One page at a time, by hand One page per request Bulk and scheduled runs

Why convert a page to Markdown for AI

  • Prompts and context windows

    Raw HTML spends context-window tokens on markup that carries no meaning. Clean Markdown gives a model signal instead of noise, which lowers the cost of each answer as well as improving it.

  • RAG pipelines

    Markdown preserves the heading hierarchy, which is what chunking relies on. Boilerplate left in a RAG index produces retrieval hits on text nobody meant to store.

  • Agent context

    Clean Markdown is what an AI agent needs before it can reason about a page. Links keep their destination, so citations survive the conversion.

  • Notes and archives

    Save an article as a Markdown file that opens in any editor and diffs cleanly in git, or pull a table out for analysis.

When this free tool is not enough

  • JavaScript-rendered pages. The converter reads the HTML the server sends. Content that appears after JavaScript runs comes back empty, and the tool says so.

  • One page, 100,000 characters. Each conversion reads the first 2 MB of HTML and returns up to 100,000 characters. Fair use is 30 conversions a minute.

  • Logins, paywalls and blocks. Pages behind a login or a CAPTCHA, or ones that refuse automated requests, report the status code they answered with. An anti-bot check returns a challenge page, not the article, so there is nothing to convert.

A single page conversion is a browser job; a thousand pages is an API job. Converting whole sites, rendering JavaScript first or running on a schedule is what MrScraper's Web Scraper API does.

How to convert HTML to Markdown with code

Fetch the page, drop the elements that are not content, and hand the rest to an HTML-to-Markdown library.

Python
import requests
from markdownify import markdownify

html = requests.get(
    "https://example.com/blog/post",
    headers={"User-Agent": "Mozilla/5.0"},
    timeout=10,
).text
print(markdownify(html, heading_style="ATX", strip=["script", "style", "nav"]))
JavaScript
import TurndownService from 'turndown'

const html = await (await fetch('https://example.com/blog/post')).text()
const turndown = new TurndownService({ headingStyle: 'atx', codeBlockStyle: 'fenced' })
turndown.remove(['script', 'style', 'nav', 'footer'])
console.log(turndown.turndown(html))

Website to text FAQ

How do I convert a website to text?

You can convert a website to text by pasting its URL into a website-to-text tool, which fetches the page, removes navigation, ads and scripts, and returns the readable content as plain text or Markdown. The converter on this page finds the main content region first, then offers the result to copy or download as a .md or .txt file.

How do I convert a web page to Markdown?

Paste the page URL into a converter and select Markdown as the output. Headings become # symbols, lists become bullets, and links are preserved in Markdown syntax with absolute URLs. Tables convert to pipe-delimited Markdown tables, and code blocks become fenced blocks with their whitespace intact, so the page reads the same outside the browser.

What is the difference between HTML and Markdown?

HTML is a markup language browsers render visually using tags. Markdown is a plain-text format using simple symbols like # and * for structure. Markdown is far more compact, which matters when feeding content to a language model: the same page measured on this page shrinks by about four fifths when converted.

Why convert web pages to Markdown for AI?

Raw HTML spends context-window tokens on navigation, scripts and styling that carry no meaning. Markdown keeps the heading hierarchy and text while removing that overhead, so a model receives more signal per token. The heading structure also gives a RAG pipeline natural chunk boundaries, and absolute links keep each chunk traceable to its source.

What gets removed when converting a website to text?

A website-to-text converter removes navigation menus, sidebars, footers, forms, scripts, styling, SVG graphics and iframes. It keeps headings, paragraphs, lists, tables, links and code blocks. When a page marks its content with an article or main element, only that region is read, so cookie banners and ads placed outside it are dropped too.

Can I convert a JavaScript-rendered page to text?

A basic converter reads the HTML the server sent, so content added by JavaScript after page load will be missing. Pages built with React, Vue or Angular usually need a tool that renders the page in a real browser first. When fewer than 20 words come back, this converter says so instead of showing an empty box.

Read the JavaScript crawling guide

Is there a free website to text converter?

Yes, the converter on this page turns any public URL into plain text or Markdown free, with no signup and no account required. Each conversion reads the first 2 MB of HTML from one page and returns up to 100,000 characters. The only other limit is a fair-use cap of 30 conversions a minute.

How do I extract text from a website in Python?

Fetch the page with requests, then parse it with BeautifulSoup and call get_text() for plain text. For Markdown output, pass the cleaned HTML through a library such as markdownify or html2text. Remove script, style and nav elements before converting, or their contents end up in the output alongside the article.

Does converting a website to text keep tables?

Yes, tables are converted to pipe-delimited Markdown tables, which preserves rows and columns as readable text without any HTML tags. Complex nested tables may lose some structure, because Markdown has no way to put a table inside a cell. The plain text output keeps the cell contents but drops the pipes.

Is it legal to convert a website to text?

Reading a publicly accessible page and reformatting it for personal use is generally permitted, but redistributing the content may infringe copyright and the site terms of service still apply. Converting a page does not change who owns the words in it. This is not legal advice; for a specific project, read the terms first.

Read: Is web scraping legal?

Your choices

Cookie preferences

Necessary cookies keep your selection. Optional categories are disabled until you switch them on.

Strictly necessary

Remembers your privacy selection and keeps the site working.

Always on