Implementation of a weblog extraction system with an improved template extraction technique

E CHANG

摘要

Purpose:The objectives of this study are to explore an effective technique to extract information from weblogs and develop an experimental system to extract structured information as much as possible with this technique.The system will lay a foundation for evaluation,analysis,retrieval,and utilization of the extracted information.Design/methodology/approach:An improved template extraction technique was proposed.Separate templates designed for extracting blog entry titles,posts and their comments were established,and structured information was extracted online step by step.A dozen of data items,such as the entry titles,posts and their commenters and comments,the numbers of views,and the numbers of citations were extracted from eight major Chinese blog websites,including Sina,Sohu and Bokee.Findings:Results showed that the average accuracy of the experimental extraction system reached 94.6％.Because the online and multi-threading extraction technique was adopted,the speed of extraction was improved with the average speed of 15 pages per second without considering the network delay.In addition,entries posted by Ajax technology can be extracted successfully.Research limitations:As the templates need to be established in advance,this extraction technique can be effectively applied to a limited range of blog websites.In addition,the stability of the extraction templates was affected by the source code of the blog pages.Practical implications:This paper has studied and established a blog page extraction system,which can be used to extract structured data,preserve and update the data,and facilitate the collection,study and utilization of the blog resources,especially academic blog resources.Originality/value:This modified template extraction technique outperforms the Web page downloaders and the specialized blog page downloaders with structured and comprehensive data extraction.

机译：目的：本研究的目的是探索一种有效的从Web日志中提取信息的技术，并开发一种实验系统以尽可能多地提取结构化信息。该系统将为评估，分析，检索和利用奠定基础。设计/方法/方法：提出了一种改进的模板提取技术。建立了用于提取博客条目标题，帖子及其评论的单独模板，并逐步在线提取了结构化信息。十几个数据项从新浪，搜狐，博基等8个主要中文博客网站中提取了条目标题，帖子及其评论和评论，观看次数和被引用次数。结果：结果显示，实验提取率达到94.6％。由于采用在线多线程提取技术，提取速度快。无需考虑网络延迟就可以以平均每秒15页的速度进行改进。此外，可以成功提取Ajax技术发布的条目。研究局限性：由于需要预先建立模板，因此该提取技术可以有效地应用有限的博客网站。此外，提取模板的稳定性还受到博客页面源代码的影响。实际意义：本文研究并建立了博客页面提取系统，该系统可用于提取结构化网页数据，保存和更新数据，并方便博客资源（尤其是学术博客资源）的收集，研究和利用。原始性/价值：这种经过修改的模板提取技术在结构上和功能上都超过了网页下载器和专业博客页面下载器。数据提取。

Implementation of a weblog extraction system with an improved template extraction technique

摘要

著录项

相关主题

期刊订阅