Copy part of any web page and paste it below. A references or citations section works best: links, ISBNs and DOIs are what the spider likes to tear off. Links and basic formatting come along with the paste.
An article about programs that walk the web, for a program that walks on web pages.
A web crawler, also called a spider or bot, is a program that browses the World Wide Web methodically, usually to build an index for a search engine. It starts from a list of seed URLs, fetches each page, pulls out its links, and adds the new ones to a queue called the crawl frontier.[1] The frontier is the crawler's to-do list, and how it is ordered decides what the crawler sees first.[2]
In June 1993 Matthew Gray wrote the World Wide Web Wanderer to measure how fast the web was growing.[3] WebCrawler, released by Brian Pinkerton at the University of Washington on 20 April 1994, was the first search engine to index the full text of every page it visited.[4] That same year Martijn Koster proposed robots.txt, a plain text file that tells crawlers which paths to leave alone.[5] It stayed an informal convention for 28 years, until the IETF published it as RFC 9309 in September 2022.[6]
The crawler behind Google was described in Brin and Page's 1998 paper, which reported fetching about 100 pages per second across several machines.[7] Mercator, built at Compaq's Systems Research Center, showed in 1999 how to make a crawler modular and extensible in Java.[8] Later systems such as UbiCrawler and IRLbot pushed past 6,000,000,000 pages.[9][10]
No crawler can fetch the whole web, so each one ranks its frontier. Najork and Wiener found that a plain breadth-first order already reaches high-quality pages early.[11]
Pages change after they are fetched. Cho and Garcia-Molina showed that revisiting every page at the same rate beats revisiting fast-changing pages more often, a result that surprised many people at the time.[12]
A crawler can fetch far faster than a small server can answer. Polite crawlers keep one connection per host, wait between requests, and honor Crawl-delay where it is set.[1]
Some sites generate endless URLs, such as a calendar with a "next month" link that never runs out. Crawlers cap depth per host and normalize URLs so they do not walk forever.[13]
The web metaphor came first, and the crawler that walks along its threads followed naturally. Real spiders walk with an alternating tetrapod gait: the first and third legs on one side move together with the second and fourth on the other, while the remaining four legs hold on.[14] The spider on this page walks the same way, and every step it takes off a link is a chance to take the link with it.