首页> 中文期刊> 《计算机应用研究》 >LD A模型下不同分词方法对文本分类性能的影响研究

LD A模型下不同分词方法对文本分类性能的影响研究

         

摘要

通过定义类别聚类密度、类别复杂度以及类别清晰度三个指标,从语料库信息度量的角度研究多种代表性的中文分词方法在隐含概率主题模型LDA下对文本分类性能的影响,定量、定性地分析不同分词方法在网页和学术文献等不同类型文本的语料上进行分类的适用性及影响分类性能的原因。结果表明:三项指标可以有效指明分词方法对语料在分类时产生的影响,Ik Analyzer和ICTCLAS分词法分别受类别复杂度和类别聚类密度的影响较大,二元分词法受三个指标的作用相当,使其对于不同语料具有较好的适应性。对于学术文献类型的语料,使用二元分词法时的分类效果较好,F1值均在80%以上;而网页类型的语料对于各种分词法的适应性更强。尝试通过对语料进行信息度量而非单纯的实验来选择提高该语料分类性能的最佳分词方法,以期为网页和学术文献等不同类型的文本在基于LDA模型的分类系统中选择合适的中文分词方法提供参考。%From the perspective of corpus measure,which includes three indicators:the clustering density,the complexity and definition of category,this paper studied the influence of three representative Chinese word segmentation methods,including IC-TCLAS,Ik Analyzer and 2-gram,on the performance of text classification under the implicit probabilistic topic model LDA.Mo-reover,the applicability of different Chinese word segmentation methods in different types of texts such as Web and academic documents and its cause were analyzed qualitatively and quantitatively.Experiments show that three indexes can effectively in-dicate the influence of word segmentation method on the classification of texts:Ik Analyzer and ICTCLAS segmentation method are more influenced respectively by the complexity of the category and the clustering density of the category,for 2-gram,the in-fluences of three indexes are similar,so it has good adaptability for different corpus.For corpus of academic literature,2-gram has better performance,F1 values are above 80%.And the corpus of Web pages is more adaptive to different word segmentation methods.This paper provides a reference for the selection of appropriate Chinese word segmentation method in classification system based on LDA model for different types of texts such as Web pages and academic literature by means of corpus measure instead of by experiments only.

著录项

相似文献

  • 中文文献
  • 外文文献
  • 专利
获取原文

客服邮箱:kefu@zhangqiaokeyan.com

京公网安备:11010802029741号 ICP备案号:京ICP备15016152号-6 六维联合信息科技 (北京) 有限公司©版权所有
  • 客服微信

  • 服务号