首页> 中文期刊> 《工程数学学报》 >粗略不相似度量及其在层次聚类中的应用

粗略不相似度量及其在层次聚类中的应用

         

摘要

局部结构特征在数据分析过程中具有重要的作用.为获得简单有效的数据集局部结构化特征检测方法,本文结合重采样误差分析和传统的近邻选择方法提出了一种检测局部结构特征的方向一致性度量—粗略不相似性度量.该度量是一种优化的近邻选择方法,不仅考虑了传统的欧氏距离排序,而且考虑了局部方向结构特征.因其计算和存储复杂度小以及具有优越的结构检测性能,可应用于无监督学习形成一种层次化的子图聚类算法—RDClust,与经典聚类算法相比,其优势在于:一是计算复杂度较小,是近似线性算法;二是无需对类的形状和分布形式做任何的假设,可自动体现数据集的局部结构;三是有一个近邻参数,且该参数对结果较鲁棒.在人工和真实数据集上的实验显示了新的度量方式应用于新算法的优越性能.%Local structural feature is important in data analysis procedure. In order to ob-tain a simple and effective feature detection method for data set's local structures, this paper proposed for detecting local structure a direction consistence measurement, rough dissimilarity, by combing re-sampling and a classical neighborhood selection method. This measurement is a optimized selection method for neighborhood, which considers not only the classical sorting method based on Euclidean distance but also the local structures of the data set. The new dissimilarity measurement can be used in unsupervised learning to construct a hierarchical sub-graph clustering, RDClust, because of the advantages of a low computation load and a good direction structure detection performance. The new clustering based on direction consistence measurement has three advantages: 1) It has a low computation load and is an approximately linear method; 2) It needs no assumption for the shape and the distribution of cluster, and can detect local structures of a data set automatically; 3) It has only one parameter which is relatively robust to clustering results. The new clustering based on direction consistence dissimilarity has good performance in testing with synthetic and real data sets.

著录项

相似文献

  • 中文文献
  • 外文文献
  • 专利
获取原文

客服邮箱:kefu@zhangqiaokeyan.com

京公网安备:11010802029741号 ICP备案号:京ICP备15016152号-6 六维联合信息科技 (北京) 有限公司©版权所有
  • 客服微信

  • 服务号