Most search engines work by crawling the Web, indexing and filtering the content they find into massive databases, and searching these databases to find results matching a particular search query.
大多数搜索引擎的工作方式是:在Web上爬行,索引并过滤在海量数据库中找到的内容,然后搜索这些数据库以找到匹配特定搜索查询的结果。
But all the major and legitimate Web crawling engines obey the requests in robots.txt.
不过,所有主要的合法Web爬虫引擎都会遵从robots . txt内的要求。
This example demonstrates the crawling phase of a Web spider.
这个例子展示了Web spider爬行的阶段。
应用推荐