Home › Glossary › What Is Web Crawling?

What Is Web Crawling?

Web crawling is the automated process of systematically browsing the web, following hyperlinks from page to page, typically to discover and index content.

Crawling vs. scraping

Crawling and scraping are often confused. A crawler’s job is discovery: it follows links across a site or the wider web to build a map of available pages. Scraping is extraction: once a page is found, a scraper pulls specific data out of it. Most large-scale data collection systems combine both — a crawler locates the pages, and a scraper extracts the target fields from each one.

Where it’s used

Search engines use crawling to build their indexes. Data collection platforms use crawling to discover new product listings, job postings, or real estate listings across a site before running extraction on each one.