首页> 中文期刊> 《浙江师范大学学报(自然科学版)》 >基于半监督学习的中文多文档子主题划分

基于半监督学习的中文多文档子主题划分

         

摘要

Aimed to depart the sub-topic of multi-documents more effectively, it was proposed a new method based on semi-supervised learning: it firstly got the primal sets of topics by hierarchy clustering based on semantic distance of sentences, and labeled the sentences which had high scores in the topics, then used the method of constrained-A>means to decide the number of topics k, and finally obtained the topic sets by k-means clustering. The experiment results indicated that this method improved the accuracy of sub-topic recognition.%为了能在多文档自动摘要过程中更好地划分子主题,提出了一种基于半监督学习的子主题划分方法:首先计算句子的语义相似度;然后通过层次聚类对可信度高的句子进行主题类别标记,生成少量已标记主题类别的句子集,在此基础上对所有句子进行constrained-k-means聚类,通过交叉验证的方法确定子主题的数目k;最后使用k-means聚类获得多文档的各个子主题.实验结果表明,该方法有效地提高了子主题的识别率.

著录项

相似文献

  • 中文文献
  • 外文文献
  • 专利
获取原文

客服邮箱:kefu@zhangqiaokeyan.com

京公网安备:11010802029741号 ICP备案号:京ICP备15016152号-6 六维联合信息科技 (北京) 有限公司©版权所有
  • 客服微信

  • 服务号