Crawling ranked deep Web data sources

详细信息查看全文

作者：Yan Wang ; Jianguo Lu ; Jessica Chen ; Yaxin Li
关键词：Deep Web crawling ; Query selection ; Estimation ; Document frequency ; Return limit
刊名：World Wide Web
出版年：2017
出版时间：January 2017
年：2017
卷：20
期：1
页码：89-110
全文大小：
刊物类别：Computer Science
刊物主题：Information Systems Applications (incl.Internet); Database Management; Operating Systems;
出版者：Springer US
ISSN：1573-1413
卷排序：20

文摘

In the era of big data, the vast majority of the data are not from the surface Web, the Web that is interconnected by hyperlinks and indexed by most general purpose search engines. Instead, the trove of valuable data often reside in the deep Web, the Web that is hidden behind query interfaces. Since numerous applications, like data integration and vertical portals, require deep Web data, various crawling methods were developed for exhaustively harvesting a deep Web data source with the minimal (or near-minimal) cost. Most existing crawling methods assume that all the documents matched by queries are returned. In practice, data sources often return the top k matches. This makes exhaustive data harvesting difficult: highly ranked documents will be returned multiple times, while documents ranked low have small chance being returned. In this paper, we decompose this problem into two orthogonal sub-problems, i.e., query and ranking bias problems, and propose a document frequency based crawling method to overcome the ranking bias problem. The rational of our method is to use the queries whose document frequencies are within the specified range to avoid the effect of search ranking plus return limit and significantly reduce the difficulty of crawling ranked data source. The method is extensively tested on a variety of datasets and compared with two existing methods. The experimental result demonstrates that our method outperforms the two algorithms by 58 % and 90 % on average respectively.

地址：北京市海淀区学院路29号邮编：100083

电话：办公室：(+86 10)66554848；文献借阅、咨询服务、科技查新：66554700