Skip to content
How to Build a Fast Web Scraper with Go
Article

How to Build a Fast Web Scraper with Go

Web Scraping

Learn how to build a fast web scraper with Go using net/http, goquery, and Colly, with guidance for concurrency, dynamic content, and managed scraping.

By MrScraper Team 7 min read

A fast web scraper in Go can start with net/http and HTML parsing. Then use goquery for selectors. Or use Colly for crawling, concurrency, and request controls. For dynamic or protected content, consider browser automation or a managed scraping service.

What Makes Go Suitable for Web Scraping?

Go has several characteristics that make it well suited for web scraping:

  • A fast compiler and lightweight runtime for efficient execution at scale
  • Strong concurrency support with goroutines and channels
  • A solid standard library for HTTP requests and HTML parsing
  • A growing ecosystem of third-party scraping libraries
  • Built-in tooling and dependency management with go mod

Building a Basic Scraper Using Go’s Standard Library

You don’t need external tools to start scraping. Go’s standard library lets you fetch and parse HTML with minimal dependencies.

Step 1: Fetching a Web Page

Use Go’s net/http package to request HTML content:

go
package main

import (
    "fmt"
    "net/http"
    "io"
)

func main() {
    resp, err := http.Get("https://example.com")
    if err != nil {
        fmt.Println("Request failed:", err)
        return
    }
    defer resp.Body.Close()

    body, _ := io.ReadAll(resp.Body)
    fmt.Println(string(body))
}

This code sends an HTTP GET request and prints the response body as raw HTML.

To extract links with a fast web scraper, use the golang.org/x/net/html package. The recursive helper walks the HTML node tree and collects href attributes. The example checks request and parsing errors before printing results.

go
package main

import (
	"fmt"
	"net/http"

	"golang.org/x/net/html"
)

func findLinks(n *html.Node) []string {
	links := []string{}
	if n.Type == html.ElementNode && n.Data == "a" {
		for _, attr := range n.Attr {
			if attr.Key == "href" {
				links = append(links, attr.Val)
			}
		}
	}
	for c := n.FirstChild; c != nil; c = c.NextSibling {
		links = append(links, findLinks(c)...)
	}
	return links
}

func main() {
	resp, err := http.Get("https://example.com")
	if err != nil {
		panic(err)
	}
	defer resp.Body.Close()
	if resp.StatusCode < 200 || resp.StatusCode >= 300 {
		panic(fmt.Errorf("unexpected HTTP status: %s", resp.Status))
	}
	root, err := html.Parse(resp.Body)
	if err != nil {
		panic(err)
	}
	for _, link := range findLinks(root) {
		fmt.Println(link)
	}
}

Leveraging Third-Party Tools: goquery

While native parsing works, third-party libraries like goquery offer a jQuery-like API that makes extraction easier.

Installing goquery

Initialize your module and install goquery:

bash
go mod init my-scraper
go get github.com/PuerkitoBio/goquery

Using goquery to Extract Data

This example prints all <h1> text from a page:

go
package main

import (
    "fmt"
    "net/http"
    "github.com/PuerkitoBio/goquery"
)

func main() {
    res, err := http.Get("https://example.com")
    if err != nil {
        fmt.Println("Request error:", err)
        return
    }
    defer res.Body.Close()

    doc, err := goquery.NewDocumentFromReader(res.Body)
    if err != nil {
        fmt.Println("Parsing error:", err)
        return
    }

    doc.Find("h1").Each(func(i int, s *goquery.Selection) {
        fmt.Println("Header:", s.Text())
    })
}

goquery wraps Go’s HTML parser with CSS selectors, making extraction more readable and maintainable.

Building Advanced Scrapers With Colly

For more complex tasks, such as crawling many pages, managing sessions, and controlling requests, use Colly. It is a popular Go framework.

Installing Colly

bash
go get -u github.com/gocolly/colly/...

Basic Colly Example

This script scrapes all links from a Wikipedia page section:

go
package main

import (
    "fmt"
    "github.com/gocolly/colly"
)

func main() {
    c := colly.NewCollector()

    c.OnHTML(".mw-parser-output", func(e *colly.HTMLElement) {
        links := e.ChildAttrs("a", "href")
        fmt.Println(links)
    })

    c.Visit("https://en.wikipedia.org/wiki/Web_scraping")
}

What this does:

  • Creates a new Colly collector
  • Registers an HTML callback using a CSS selector
  • Visits the page and extracts all matching links

Colly reduces boilerplate and supports rate limiting, cookies, retries, and more.

Performance and Concurrency

Go is well suited to building a fast web scraper across many pages. Goroutines enable parallel HTTP requests, while channels coordinate data flow safely without manual thread management. Add appropriate rate limiting and synchronization so the scraper can process high-throughput workloads efficiently.

Fast Web Scraper Worker Pool

A fast web scraper uses a bounded worker pool. Goroutines fetch jobs at the same time. A sync.WaitGroup lets the program close results after all workers exit.

go
package main

import (
	"fmt"
	"io"
	"net/http"
	"sync"
)

type result struct {
	url  string
	code int
	err  error
}

func worker(client *http.Client, jobs <-chan string, results chan<- result, wg* sync.WaitGroup) {
	defer wg.Done()
	for url := range jobs {
		resp, err := client.Get(url)
		if err != nil {
			results <- result{url: url, err: err}
			continue
		}
		_, readErr := io.Copy(io.Discard, resp.Body)
		resp.Body.Close()
		if readErr != nil {
			results <- result{url: url, code: resp.StatusCode, err: readErr}
			continue
		}
		results <- result{url: url, code: resp.StatusCode}
	}
}

func main() {
	urls := []string{
		"https://example.com",
		"https://example.org",
		"https://example.net",
	}
	client := &http.Client{}
	jobs := make(chan string)
	results := make(chan result)
	var wg sync.WaitGroup

	const workers = 3
	wg.Add(workers)
	for i := 0; i < workers; i++ {
		go worker(client, jobs, results, &wg)
	}

	go func() {
		for _, url := range urls {
			jobs <- url
		}
		close(jobs)
		wg.Wait()
		close(results)
	}()

	for item := range results {
		if item.err != nil {
			fmt.Printf("%s: %v\n", item.url, item.err)
			continue
		}
		fmt.Printf("%s: HTTP %d\n", item.url, item.code)
	}
}

The fixed worker count limits in-flight requests. Closing the results channel after Wait prevents lost output. It also avoids a range loop that never ends. Add parsing inside each worker, then introduce request pacing and retries that match the target site's rules.

Handling Dynamic or Protected Content

Go-based scrapers work best for static content, but modern sites often rely on JavaScript or protected APIs. In such cases, you can:

  • Reverse-engineer underlying API requests instead of scraping HTML
  • Use browser automation tools like chromedp or headless Chrome bindings
  • Combine Go scrapers with rendering services for JavaScript-heavy pages

Each approach involves trade-offs between performance and complexity.

MrScraper: A Managed Scraping Solution

For teams that want to avoid maintaining scraping infrastructure, managed solutions can simplify development.

MrScraper provides:

  • Automatic proxy rotation and anti-bot handling
  • JavaScript rendering support
  • Clean, structured JSON output
  • API-based scraping jobs that integrate easily with Go applications

Instead of managing retries, IPs, and browser automation yourself, call MrScraper’s API. Focus on using the data.

Conclusion

Web scraping with Go offers a strong balance of performance and flexibility. You can start with Go’s standard libraries for simple scrapers. Use goquery for easier DOM traversal. Scale up with Colly for advanced crawling tasks. Go’s concurrency model works well for high-throughput scraping. When you need JavaScript rendering or must handle anti-bot measures, use browser automation. Managed scraping services can also fill the gap. With the right tools and patterns, Go enables reliable and scalable data extraction for modern applications.

What We Learned

A fast web scraper is not defined by request speed alone. The durable pattern is a bounded pipeline with explicit extraction rules, respectful request limits, and recoverable progress.

Use a checkpoint after each completed page so a stopped run resumes without repeating successful work. Keep fetching, parsing, and persistence separate, then test each stage independently.

go
package main

import (
	"encoding/json"
	"fmt"
	"os"
	"path/filepath"
)

type Checkpoint struct {
	Done map[string]bool `json:"done"`
}

func save(path string, state Checkpoint) error {
	tmp, err := os.CreateTemp(filepath.Dir(path), "checkpoint-*.tmp")
	if err != nil {
		return err
	}
	tmpName := tmp.Name()
	defer os.Remove(tmpName)

	if err := json.NewEncoder(tmp).Encode(state); err != nil {
		tmp.Close()
		return err
	}
	if err := tmp.Close(); err != nil {
		return err
	}
	return os.Rename(tmpName, path)
}

func main() {
	state := Checkpoint{Done: map[string]bool{}}
	for _, page := range []string{"page-1", "page-2", "page-3"} {
		if state.Done[page] {
			continue
		}
		fmt.Println("process", page)
		state.Done[page] = true
		if err := save("checkpoint.json", state); err != nil {
			panic(err)
		}
	}
}
  • Start with the standard library when request and parsing needs are simple.
  • Use goquery or Colly when selectors, crawling, and request coordination would otherwise become repetitive.
  • Bound concurrency and honor the target site’s access rules instead of treating goroutines as unlimited capacity.
  • Add checkpoints, timeouts, retries, and observable errors before scaling a scraper beyond a one-off run.
  • Switch to an API, browser automation, or a managed service when the target needs rendering or protection. HTML requests cannot provide these features.

Start Planning Your Go Scraping Workflow

Use this quickstart guide to choose a Go scraping approach and evaluate where a managed workflow may fit your project.

Get Started

Summarize this post

Open it in your assistant of choice with the prompt ready to send.

Take a Taste of Easy Scraping!

Your choices

Cookie preferences

Necessary cookies keep your selection. Optional categories are disabled until you switch them on.

Strictly necessary

Remembers your privacy selection and keeps the site working.

Always on