...
首页> 外文期刊>Modern Applied Science >Ontology Based Fuzzy Document Clustering Scheme
【24h】

Ontology Based Fuzzy Document Clustering Scheme

机译:基于本体的模糊文档聚类方案

获取原文

摘要

Document clustering is the technique used to group up the document with the reference to the similarity. It is widely used in web mining and digital library environment. Documents are represented in vector space model. Each document is a vector in the word space and each element of the vector indicates the frequency of the corresponding word in the document. Documents are presented as high dimensional data elements. It is a very complex task to cluster documents using K-means clustering algorithm. The sub space clustering schemes can be adopted to cluster documents. The document clustering uses the term weights from the similarity measure. The sub space model uses the relevant attributes for the similarity estimation. The fuzzy logic is used to cluster the documents. The fuzzy document clustering scheme is enhanced with semantic analysis mechanism. Semantic analysis is carried out with the support of the ontology. The ontology is used to maintain term relationships. Term relationships are represented using the synonym, meronym and hypernym factors. Ontology is manually collected by the users. Domain based ontology is used for the document clustering process. The system uses the data mining domain based ontology for the semantic analysis. Semantic weights are used in the similarity measure. Fuzzy based text document clustering scheme uses the stop word filters and stemming process under the document preprocess. Term clustering and semantic clustering operations are performed in the system.
机译:文档聚类是用于参考相似性对文档进行分组的技术。它广泛用于Web挖掘和数字图书馆环境。文档在向量空间模型中表示。每个文档都是单词空间中的一个向量,该向量的每个元素都表示文档中相应单词的频率。文档以高维数据元素的形式呈现。使用K-means聚类算法对文档进行聚类是一项非常复杂的任务。可以采用子空间聚类方案对文档进行聚类。文档聚类使用相似性度量中的术语权重。子空间模型使用相关属性进行相似性估计。模糊逻辑用于对文档进行聚类。通过语义分析机制增强了模糊文档聚类方案。语义分析是在本体的支持下进行的。本体用于维护术语关系。术语关系使用同义词,同义词和上位词因子表示。本体由用户手动收集。基于域的本体用于文档聚类过程。该系统使用基于数据挖掘域的本体进行语义分析。语义权重用于相似性度量中。基于模糊的文本文档聚类方案在文档预处理过程中使用了停用词过滤器和词干处理功能。术语聚类和语义聚类操作在系统中执行。

著录项

相似文献

  • 外文文献
  • 中文文献
  • 专利
获取原文

客服邮箱:kefu@zhangqiaokeyan.com

京公网安备:11010802029741号 ICP备案号:京ICP备15016152号-6 六维联合信息科技 (北京) 有限公司©版权所有
  • 客服微信

  • 服务号