首页> 中文期刊> 《计算机工程与应用》 >频繁项集在Deep Web数据源聚类中的应用

频繁项集在Deep Web数据源聚类中的应用

         

摘要

在Deep Web页面的背后隐藏着海量的可以通过结构化的查询接口进行访问的数据源.将这些数据源按所属领域进行组织划分,是Deep Web数据集成中的一个关键步骤.已有的划分方法主要是基于查询接口模式和提交查询返回结果,存在查询接口特征难以完全抽取和提交数据库查询效率不高等问题.提出了一种结合网页文本信息,基于频繁项集的聚类方法,根据数据源查询接口所在页面的标题、关键词和提示文本,将数据源按照领域进行聚类,有效解决了传统方法中依赖查询接口特征以及文本模型的高维性问题.实验结果表明该方法是可行的,具有较高的效率.%There are thousands of data sources hiding behind the Deep Web pages which can be accessed through structured query interfaces. Organizing these data sources by their domains has become an important step in Deep Web data integration process. Existing methods mainly focus on query interface schema and query results which have the disadvantages of difficulty in extracting interface schemas and deficiency of submitting queries to the database. A method based on frequent itemsets is proposed to cluster the data sources by their domains. This method considers the Web page text information such as title, key words and label text and solves the problems of overdepen-dency on the query interface and high dimensionality of text processing in traditional solutions. Experimental results show effectiveness and efficiency of this method.

著录项

相似文献

  • 中文文献
  • 外文文献
  • 专利
获取原文

客服邮箱:kefu@zhangqiaokeyan.com

京公网安备:11010802029741号 ICP备案号:京ICP备15016152号-6 六维联合信息科技 (北京) 有限公司©版权所有
  • 客服微信

  • 服务号