基于网页结构相似度的Web 信息抽取

聂卉

首页> 中文期刊> 《情报学报》 >基于网页结构相似度的Web 信息抽取

基于网页结构相似度的Web 信息抽取

开具论文收录证明 >>

期刊封面封底目录下载 >>

文献代查 >>

页面导航

摘要
著录项
相似文献
相关主题

摘要

In this paper, we focus on Web Information Extraction by use of structural similarities.Combined with URL similarity, the Tree- Edit- Distance method is adapted to measure the similarity between web pages.The similarity is used for detecting changes between the pages as the criterion.By this way, the suitable extraction rules will be selected, otherwise,the target page will be fed to rule learner module.Extraction rules are induced automatically by machine learning.The introduction of the algorithm given extraction system the self-renewal capacity , which make it better adapt to the practical application for the extraction of heterogeneous resources.The feasibility and efficiency of the algorithm in the prototype system was able to verify.%本文重点探讨基于编辑距离的网页相似度算法在Web 抽取系统中的应用与实现.通过结合基于URL 及编辑距离的网页结构相似度的计算方法,抽取系统在抽取过程中能够检测网页结构的变化,从而主动做出判断,选择适应规则进行抽取或通过主动学习自动扩展规则库.结构相似度计算赋予系统感知网页结构变化的能力,系统通过主动自我更新与调整,能更好地适应面向实际应用的异构资源的获取.算法的可行性和效率在原型系统中得以验证.

著录项

来源
《情报学报》 |2011年第3期|268-274|共7页
作者
聂卉;
展开▼
作者单位

中山大学资讯管理系;

广州;

510275;

展开▼
原文格式 PDF
正文语种 chi
中图分类
关键词
Web; 信息抽取; 结构相似度; 编辑距离;

相似文献

中文文献
外文文献
专利

1. 基于领域本体的Web信息抽取方法的设计与实现——以网易汽车资讯网页信息抽取为例 [J] . 吴恒亮 . 图书馆论坛 . 2010,第003期
2. 基于WEB网页文本信息抽取研究与实现 [J] . 刘三星1 . 数据挖掘 . 2015,第004期
3. 基于网页结构的WEB信息抽取系统设计 [J] . 张彩月 . 计算机光盘软件与应用 . 2012,第006期
4. 基于网页聚类的Web信息自动抽取 [J] . 邱韬奋 ,杨天奇 ,曾洪波 . 微型机与应用 . 2011,第004期
5. 基于DOM和网页模板的Web信息抽取 [J] . 王丽 ,唐建雄 . 电脑知识与技术 . 2007,第018期
6. 基于统计的中文网页正文信息抽取方法研究 [C] . 李芳芳 ,葛斌 . 第三届全国社会计算会议、平行控制会议、平行管理会议 . 2011
7. 基于模板化网络爬虫技术的Web网页信息抽取 [A] . 乔峰 . 2012

基于网页结构相似度的Web 信息抽取

摘要

著录项

相似文献

相关主题

期刊订阅