首页> 中文期刊>中文信息学报 >基于广义话题理论的话题句识别

基于广义话题理论的话题句识别

     

摘要

Nowadays the Chinese machine translation and information extraction is still far from satisfactory. One important reason is that the topics are often omitted in the head of Chinese Punctuation Clause (abbreviated as PClause). Based on the Generalized Topic Theory, this paper proposes a novel method for topic clause identification from PClause based on the characteristic of topic strcture. The method consists of two tasks in practice: topic clause identification from a single PClause and topic clause construction for a series of PClauses. In the first task,semantic generalization and edit distance are applied in this paper, and the accuracy rate for open test is 12. 51% higher than baseline. The result proves the effectiveness of the generalized topic theory in topic clause identification from a single PClause.%汉语标点句句首话题缺失是机器翻译、信息抽取准确率不高的原因之一.该文从广义话题理论出发,根据汉语话题结构的特点,提出标点句的话题句识别研究方案,包括两个阶段性任务:单个标点句的话题句识别和序列标点句的话题句序列构建.识别出标点句的话题句也就找到了标点句句首缺失的话题.该文解决单个标点句的话题句识别任务,主要采用语义泛化和编辑距离两种手段.实验中开放测试的准确率比基线高出12.51个百分点.该结果说明,运用广义话题理论进行单个标点句的话题句识别可产生明显的效果.

著录项

相似文献

  • 中文文献
  • 外文文献
  • 专利
获取原文

客服邮箱:kefu@zhangqiaokeyan.com

京公网安备:11010802029741号 ICP备案号:京ICP备15016152号-6 六维联合信息科技 (北京) 有限公司©版权所有
  • 客服微信

  • 服务号