基于语义相关度主题爬虫的语料采集方法

英文篇名：Corpus Collection Based on Semantic Relevancy Focused Crawler
作者：周昆 ; 王钊 ; 于碧辉
英文作者：ZHOU Kun;WANG Zhao;YU Bi-Hui;University of Chinese Academy of Sciences;Shenyang Institute of Computing Technology, Chinese Academy of Sciences;Center for Information Technology,Shenyang State Tax Bureau;
关键词：生语料采集 ; 语义相关度主题爬虫 ; 页面信息相关度 ; URL结构信息
英文关键词：corpus collection;;semantic relevancy focused crawler;;page information semantic relevancy;;URL structural information
中文刊名：XTYY
英文刊名：Computer Systems & Applications
机构：中国科学院大学;中国科学院沈阳计算技术研究所;沈阳市国家税务局信息中心;
出版日期：2019-05-15
出版单位：计算机系统应用
年：2019
期：v.28
语种：中文;
页：XTYY201905029
页数：6
CN：05
ISSN：11-2854/TP
分类号：192-197

摘要

针对特定领域语料采集任务,设计了基于语义相关度主题爬虫的语料采集方法.根据选定的主题词,利用页面描述信息,基于维基百科中文语料训练出的词分布式表示综合HowNet计算页面信息相关度,结合URL的结构信息预测未访问URL链指的页面内容与特定领域的相关程度.实验表明,系统能够有效的采集互联网中的党建领域页面内容作为党建领域生语料,在党建领域网站上的平均准确率达到94.87%,在门户网站上的平均准确率达到64.20%.
To address the corpus collection, the corpus collection system based on semantic relevancy focused crawler is implemented. Word vector trained by Wikipedia and HowNet are used for calculating page information semantic relevancy with descriptive information according to topical keywords, and the URL structural information is used for calculating the topical relevancy. Experimental results show that this system has better effect on party-construction corpus collection with high precision of average accurate rate 94.87%, while the average accurate rate for web pages is 64.20%.

引文

1Hersovici M,Jacovi M,Maarek YS,et al.The shark-search algorithm.An application:Tailored Web site mapping.Proceedings of the 7th International Conference on World Wide Web.Brisbane,Australia.1998.317-326.
    2苏祺,项锟,孙斌.基于链接聚类的Shark-Search算法.第四届全国搜索引擎和网上信息挖掘学术研讨会论文集.济南.2006.
    3黄仁,王良伟.基于主题相关概念和网页分块的主题爬虫研究.计算机应用研究,2013,30(8):2377-2380,2409.[doi:10.3969/j.issn.1001-3695.2013.08.034]
    4魏勇,胡丹露,郝晨光,等.基于分类关键词词频模型的地缘政治主题爬虫设计.计算机工程,2016,42(2):45-50.[doi:10.3969/j.issn.1000-3428.2016.02.008]
    5Almpanidis G,Kotropoulos C,Pitas I.Combining text and link analysis for focused crawling-An application for vertical search engines.Information Systems,2007,32(6):886-908.[doi:10.1016/j.is.2006.09.004]
    6Du YJ,Liu WJ,Lv XJ,et al.An improved focused crawler based on semantic similarity vector space model.Applied Soft Computing,2015,36:392-407.[doi:10.1016/j.asoc.2015.07.026]
    7Suebchua T,Manaskasemsak B,Rungsawang A,et al.Efficient topical focused crawling through neighborhood feature.New Generation Computing,2018,36(2):95-118.[doi:10.1007/s00354-017-0029-8]
    8刘群,李素建.基于《知网》的词汇语义相似度计算.中文计算语言学,2002,7(2):59-76.
    9董振东.HowNet.http://www.keenage.com,2013.

地址：北京市海淀区学院路29号邮编：100083

电话：办公室：(+86 10)66554848；文献借阅、咨询服务、科技查新：66554700