Pattern-Based Automatic Parallelization of Representative-Based Clustering Algorithms

机译：基于模式的基于模式的自动并行化的代表性聚类算法

获取原文

页面导航

摘要
著录项
相似文献
相关主题

摘要

Ease of programming and optimal parallel performance have historically been on the opposite side of a tradeoff, forcing the user to choose. With the advent of the Big Data era and rapid evolution of sequential algorithms, the data analytics community can no longer afford the tradeoff. We observed that several clustering algorithms often share common traits - particularly, algorithms belonging to same class of clustering exhibit significant overlap in processing steps. Here, we present our observation on domain patterns in Representative-based clustering algorithms and how they manifest as clearly identifiable programming patterns when mapped to a Domain Specific Language (DSL). We have integrated the signatures of these patterns in the DSL compiler for parallelism identification and automatic parallel code generation. Our experiments on different state-of-the-art parallelization frameworks shows that our system is able to achieve near-optimal speedup while requiring a fraction of the programming effort, making it an ideal choice for the data analytics community.

机译：历史上，易于编程和最佳平行性能，迫使用户选择。随着序贯算法的大数据时代和快速演变的出现，数据分析社区无法再提供权衡。我们观察到，几种聚类算法通常共享常见的特征 - 特别地，属于同一类聚类的算法在处理步骤中表现出显着的重叠。在这里，我们在基于代表性的聚类算法中的域模式的观察以及它们在映射到域特定语言（DSL）时如何表现为清晰可识别的编程模式。我们在DSL编译器中集成了这些模式的签名，以进行并行识别和自动并行代码生成。我们对不同最先进的并行化框架的实验表明，我们的系统能够在需要一小部分编程工作的同时实现近最佳的加速，使其成为数据分析社区的理想选择。

著录项

来源
《IEEE International Conference on Data Science and Advanced Analytics》|2019年|696p|共10页
会议地点
作者
Saiyedul Islam; Sundar Balasubramaniam; Shruti Gupta; Shikhar Brajesh; Rohan Badlani; Nitin Labhishetty; Abhinav Baid; Poonam Goyal; Navneet Goyal;
展开▼
作者单位

展开▼
会议组织
原文格式 PDF
正文语种
中图分类总体结构、系统结构;
关键词
Kernel; Clustering algorithms; Programming; Data analysis; DSL; C++ languages; Program processors;

机译：内核;聚类算法;编程;数据分析;DSL;C ++语言;程序处理器;

相似文献

外文文献
中文文献
专利

1. Automatic parallelization of representative-based clustering algorithms for multicore cluster systems [J] . Saiyedul Islam, Sundar Balasubramaniam, Shruti Gupta, International Journal of Data Science and Analytics . 2020,第2期

机译：用于多核群集系统的基于代表性聚类算法的自动并行化
2. An architecture for component-based design of representative-based clustering algorithms [J] . Boris Delibasic, Milan Vukicevic, Milos Jovanovic, Data & Knowledge Engineering . 2012,第期

机译：基于代表的聚类算法的基于组件设计的体系结构
3. A design pattern-based approach for automatic choice of semi-partitioned and global scheduling algorithms [J] . Magdich Amina, Kacem Yessine Hadj, Kerboeuf Mickael, Information and software technology . 2018,第MAY期

机译：基于设计模式的半分区和全局调度算法自动选择方法
4. Pattern-Based Automatic Parallelization of Representative-Based Clustering Algorithms [C] . Saiyedul Islam, Sundar Balasubramaniam, Shruti Gupta, IEEE International Conference on Data Science and Advanced Analytics . 2019

机译：基于模式的基于模式的自动并行化的代表性聚类算法
5. On density-based and representative-based spatial clustering algorithms. [D] . Chen, Chun-Sheng. 2011

机译：基于密度和基于代表的空间聚类算法。
6. Analysis of Parallel Algorithms on SMP Node and Cluster of Workstations Using Parallel Programming Models with New Tile-based Method for Large Biological Datasets [O] . D. D. Shrimankar, S. R. Sathe 2016

机译：大型生物数据集基于新图块的并行编程模型对SMP节点和工作站集群的并行算法进行分析
7. Automatic Parallelizing Compiler For Distributed Memory Parallel Computers: New Algorithms To Improve The Performance Of The Inspector/executor [O] . Atsushi Kubota, Ikuo Miyoshi, Hiroshi Nakashima, 100

机译：分布式内存并行计算机的自动并行编译器：提高检查器/执行器性能的新算法

Pattern-Based Automatic Parallelization of Representative-Based Clustering Algorithms

摘要

著录项

相似文献

相关主题

期刊订阅