Pattern-Based Automatic Parallelization of Representative-Based Clustering Algorithms

机译：基于模式的基于代表的聚类算法的自动并行化

获取原文

页面导航

摘要
著录项
引文网络
相似文献
相关主题

摘要

Ease of programming and optimal parallel performance have historically been on the opposite side of a tradeoff, forcing the user to choose. With the advent of the Big Data era and rapid evolution of sequential algorithms, the data analytics community can no longer afford the tradeoff. We observed that several clustering algorithms often share common traits - particularly, algorithms belonging to same class of clustering exhibit significant overlap in processing steps. Here, we present our observation on domain patterns in Representative-based clustering algorithms and how they manifest as clearly identifiable programming patterns when mapped to a Domain Specific Language (DSL). We have integrated the signatures of these patterns in the DSL compiler for parallelism identification and automatic parallel code generation. Our experiments on different state-of-the-art parallelization frameworks shows that our system is able to achieve near-optimal speedup while requiring a fraction of the programming effort, making it an ideal choice for the data analytics community.

机译：一直以来，易于编程和最佳并行性能一直处于权衡的对立面，迫使用户进行选择。随着大数据时代的到来以及顺序算法的迅速发展，数据分析社区再也无法承受这种折衷。我们观察到几种聚类算法通常具有共同的特征-特别是，属于同一类聚类的算法在处理步骤中表现出明显的重叠。在这里，我们介绍对基于代表的聚类算法中的域模式的观察，以及在映射到特定域语言（DSL）时它们如何表现为可清晰识别的编程模式。我们已将这些模式的签名集成到DSL编译器中，以进行并行性识别和自动并行代码生成。我们在不同的最新并行化框架上进行的实验表明，我们的系统能够实现近乎最佳的加速，而只需要少量的编程工作，使其成为数据分析社区的理想选择。

著录项

来源
《IEEE International Conference on Data Science and Advanced Analytics》|2018年|99-108|共10页
会议地点
作者
Saiyedul Islam; Sundar Balasubramaniam; Shruti Gupta; Shikhar Brajesh; Rohan Badlani; Nitin Labhishetty; Abhinav Baid; Poonam Goyal; Navneet Goyal;
展开▼
作者单位

展开▼
会议组织
原文格式 PDF
正文语种
中图分类
关键词
Kernel; Clustering algorithms; Programming; Data analysis; DSL; C++ languages; Program processors;

机译：内核;聚类算法;编程;数据分析; DSL; C ++语言;程序处理器;

相似文献

外文文献
中文文献
专利

1. Automatic parallelization of representative-based clustering algorithms for multicore cluster systems [J] . Saiyedul Islam, Sundar Balasubramaniam, Shruti Gupta, International Journal of Data Science and Analytics . 2020,第2期

机译：用于多核群集系统的基于代表性聚类算法的自动并行化
2. An architecture for component-based design of representative-based clustering algorithms [J] . Boris Delibasic, Milan Vukicevic, Milos Jovanovic, Data & Knowledge Engineering . 2012,第期

机译：基于代表的聚类算法的基于组件设计的体系结构
3. A design pattern-based approach for automatic choice of semi-partitioned and global scheduling algorithms [J] . Magdich Amina, Kacem Yessine Hadj, Kerboeuf Mickael, Information and software technology . 2018,第MAY期

机译：基于设计模式的半分区和全局调度算法自动选择方法
4. Pattern-Based Automatic Parallelization of Representative-Based Clustering Algorithms [C] . Saiyedul Islam, Sundar Balasubramaniam, Shruti Gupta, IEEE International Conference on Data Science and Advanced Analytics . 2019

机译：基于模式的基于模式的自动并行化的代表性聚类算法
5. On density-based and representative-based spatial clustering algorithms. [D] . Chen, Chun-Sheng. 2011

机译：基于密度和基于代表的空间聚类算法。
6. Analysis of Parallel Algorithms on SMP Node and Cluster of Workstations Using Parallel Programming Models with New Tile-based Method for Large Biological Datasets [O] . D. D. Shrimankar, S. R. Sathe 2016

机译：大型生物数据集基于新图块的并行编程模型对SMP节点和工作站集群的并行算法进行分析
7. Automatic Parallelizing Compiler For Distributed Memory Parallel Computers: New Algorithms To Improve The Performance Of The Inspector/executor [O] . Atsushi Kubota, Ikuo Miyoshi, Hiroshi Nakashima, 100

机译：分布式内存并行计算机的自动并行编译器：提高检查器/执行器性能的新算法

Pattern-Based Automatic Parallelization of Representative-Based Clustering Algorithms

摘要

著录项

引文网络

相似文献

相关主题

期刊订阅