Parallelizing dense and banded linear algebra libraries using SMPSs

Rosa M. Badia; Jose R. Herrero; Jesus Labarta; Josep M. Perez; Enrique S. Quintana-Orti; Gregorio Quintana-Orti

首页> 外文期刊>Concurrency and Computation >Parallelizing dense and banded linear algebra libraries using SMPSs

【24h】

Parallelizing dense and banded linear algebra libraries using SMPSs

机译：使用SMPS并行化稠密和带状线性代数库

获取原文

获取原文并翻译 | 示例

获取外文期刊封面封底 >>

开具论文收录证明 >>

文献代查 >>

页面导航

摘要
著录项
相似文献
相关主题

摘要

The promise of future many-core processors, with hundreds of threads running concurrently, has led the developers of linear algebra libraries to rethink their design in order to extract more parallelism, further exploit data locality, attain better load balance, and pay careful attention to the critical path of computation. In this paper we describe how existing serial libraries such as (C)LAPACK and FLAME can be easily parallelized using the SMPSs tools, consisting of a few OpenMP-like pragmas and a runtime system. In the LAPACK case, this usually requires the development of blocked algorithms for simple BLAS-level operations, which expose concurrency at a finer grain. For better performance, our experimental results indicate that column-major order, as employed by this library, needs to be abandoned in benefit of a block data layout. This will require a deeper rewrite of LAPACK or, alternatively, a dynamic conversion of the storage pattern at run-time. The parallelization of FLAME routines using SMPSs is simpler as this library includes blocked algorithms (or algorithms-by-blocks in the FLAME argot) for most operations and storage-by-blocks (or block data layout) is already in place.

机译：未来的具有数百个线程并发运行的多核处理器的承诺，导致线性代数库的开发人员重新考虑其设计，以提取更多的并行性，进一步利用数据局部性，获得更好的负载平衡并特别注意计算的关键路径。在本文中，我们描述了如何使用SMPS工具轻松地并行化现有的串行库，例如（C）LAPACK和FLAME，该工具由一些类似OpenMP的编译指示和运行时系统组成。在LAPACK情况下，这通常需要开发用于简单BLAS级操作的分块算法，从而以更精细的粒度公开并发。为了获得更好的性能，我们的实验结果表明，需要放弃此库采用的列优先顺序，以便利用块数据布局。这将需要对LAPACK进行更深的重写，或者在运行时对存储模式进行动态转换。使用SMPS进行FLAME例程的并行化更为简单，因为该库包含针对大多数操作的分块算法（或FLAME Argot中的逐块算法），并且逐块存储（或块数据布局）已经到位。

著录项

来源
《Concurrency and Computation》 |2009年第18期|2438-2456|共19页
作者
Rosa M. Badia; Jose R. Herrero; Jesus Labarta; Josep M. Perez; Enrique S. Quintana-Orti; Gregorio Quintana-Orti;
展开▼
作者单位

Barcelona Supercomputing Center, Centro Nacional de Supercomputacion (BSC-CNS) and Universitat Politecnica de Catalunya, Nexus II Building, C. Jordi Girona 29, 08034 Barcelona, Spain Consejo Superior de Investigaciones Cientificas (CSIC), Spain;

Departmento de Arquitectura de Computadores, Universitat Politecnica de Catalunya, 08034 Barcelona, Spain;

Barcelona Supercomputing Center, Centro Nacional de Supercomputacion (BSC- CNS) and Universitat Politecnica de Catalunya, Nexus II Building, C. Jordi Girona 29, 08034 Barcelona, Spain;

Barcelona Supercomputing Center, Centro Nacional de Supercomputacion (BSC- CNS) and Universitat Politecnica de Catalunya, Nexus II Building, C. Jordi Girona 29, 08034 Barcelona, Spain;

Departmento de Ingeniena y Ciencia de Computadores, Universidad Jaume I, 12.071 Castellon, Spain;

Departmento de Ingeniena y Ciencia de Computadores, Universidad Jaume I, 12.071 Castellon, Spain;

展开▼
收录信息
原文格式 PDF
正文语种 eng
中图分类
关键词
linear algebra libraries; programmability; high performance; dynamic scheduling; multi-core processors;

机译：线性代数库可编程性高性能;动态调度;多核处理器;

相似文献

外文文献
中文文献
专利

1. Leveraging task-parallelism in message-passing dense matrix factorizations using SMPSs [J] . Alberto F. Martin, Ruyman Reyes, Rosa M. Badia, Parallel Computing . 2014,第5a6期

机译：在使用SMPS的消息传递密集矩阵分解中利用任务并行性
2. Recursion based parallelization of exact dense linear algebra routines for Gaussian elimination [J] . Dumas Jean-Guillaume, Gautier Thierry, Pernet Clement, Parallel Computing . 2016,第SEPa期

机译：基于递归的精确密集线性代数例程的高斯消除并行化
3. DSSLIB(tm)– A Library of Parallelized and Optimized Linear Algebra Subroutines for SPARC Computers [J] . Gregory V.Wilson Scientific programming . 1995,第1期

机译：DSSLIB（tm）–用于SPARC计算机的并行和优化的线性代数子例程库
4. Parallelizing Dense Linear Algebra Operations with Task Queues in llc [C] . Antonio J. Dorta, Jose M. Badia, Enrique S. Quintana-Orti, European PVM/MPI Users, Group Meeting; 20070930-1003; Paris(FR) . 2007

机译：在llc中使用任务队列并行进行密集线性代数运算
5. Parallelization for MIMD multiprocessors with applications to linear algebra algorithms. [D] . Nelken, Israel Haim. 1989

机译：MIMD多处理器的并行化及其在线性代数算法中的应用。
6. Algebraic Reconstruction Technique (ART) for parallel imaging reconstruction of undersampled radial data: Application to cardiac cine [O] . Shu Li, Cheong Chan, Jason P. Stockmann, -1

机译：用于欠采样径向数据的并行成像重建的代数重建技术（ART）：在心脏电影中的应用
7. Parallelizing dense and banded linear algebra libraries using SMPSs [O] . Badía Sala Rosa María, Herrero Josep R., Labarta Mancho Jesús, 2009

机译：使用smps并行化密集和带状线性代数库
8. Design of a parallel, dense linear algebra software library: Reduction to Hessenberg, tridiagonal, and bidiagonal form [R] . Choi, J., Dongarra, J. J., Walker, D. W. 1994

机译：并行，密集线性代数软件库的设计：简化为Hessenberg，三对角和双对角形式

Parallelizing dense and banded linear algebra libraries using SMPSs

摘要

著录项

相似文献

相关主题

期刊订阅