首页> 中文期刊>计算机应用与软件 >汉维主题网页自动获取技术的研究

汉维主题网页自动获取技术的研究

     

摘要

In order to obtain a lot of Chinese-Uighur corpus for machine translation research, the paper proposes a method to automatically acquire topic information from webpages. Considering the comparatively more concenrrative topic information, higher text density and a lot of noise information brought about by links in theme web pages, the algorithm proposed in the paper first of all classifies links into noise links and non-noise links, then edit the source to remove anchor texts from noise links and HTML tags from non-noise links; next according to container tags split the source code into parts, then delete those source blocks whose text lengths and densities are all less than their respective threshold values. Experiment results on Chinese-Uighur webpages illustrate that, when configured with appropriate threshold values the algorithm's execution result reaches 90% fine rate.%为了获得大量用于机器翻译研究的汉维(维吾尔)文语料,提出一种从网页中自动获取主题信息的方法.考虑到有主题网页中主题信息分布相对集中、文本密度较高,并且这类网页中大量的噪音信息是由链接引入的,提出的算法首先将链接分为噪音链接和非噪音链接,并在源码中删除噪音链接的锚文本和非噪音链接的HTML标签,然后利用容器标签将源码划分为若干部分并删除文本长度和文本密度均小于各自阚值的源码块.针对汉维网页做了实验,实验结果表明,算法在设置合适的阈值的情况下良好率达到90%以上.

著录项

相似文献

  • 中文文献
  • 外文文献
  • 专利
获取原文

客服邮箱:kefu@zhangqiaokeyan.com

京公网安备:11010802029741号 ICP备案号:京ICP备15016152号-6 六维联合信息科技 (北京) 有限公司©版权所有
  • 客服微信

  • 服务号