命名实体识别是自然语言处理领域的一个重要任务,为许多上层应用提供支持。本文主要研究汉语开放域命名实体边界的识别。由于目前该任务尚缺乏训练语料,而人工标注语料的代价又太大,本文首先基于双语平行语料和英语句法分析器自动标注了一个汉语专有名词语料,另外基于汉语依存树库生成了一个名词复合短语语料,然后使用自学习方法将这两部分语料融合形成命名实体边界识别语料,同时训练边界识别模型。实验结果表明自学习的方法可以提高边界识别的准确率和召回率。%Named entity recognition is an important task in the domain of Natural Language Processing,which plays an im-portant role in many applications.This paper focuses on the boundary identification of Chinese open -domain named enti-ties.Because the shortage of training data and the huge cost of manual annotation,the paper proposes a self -training ap-proach to identify the boundaries of Chinese open -domain named entities in context.Due to the lack of training data,the paper firstly generates a large scale Chinese proper noun corpus based on parallel corpora,and also transforms a Chinese dependency tree bank to a noun compound training corpus.Subsequently,the paper proposes a self -training -based ap-proach to combine the two corpora and train a model to identify boundaries of named entities.The experiments show the proposed method can take full advantage of the two corpora and improve the performance of named entity boundary identifi-cation.
展开▼