论文信息 - Incorporating memory layout in the modeling of message passing programs

Incorporating memory layout in the modeling of message passing programs

One of the most fundamental tasks of an automatic parallelization tool is to find an optimal domain decomposition for a given application. For regular domain problems (such as simple matrix manipulations) this task may seem trivial. However, communication costs in message passing programs often significantly depend on the memory layout of data blocks to be transmitted. As a consequence, straightforward domain decompositions may be non-optimal. In this paper we introduce a new point-to-point communication model (called P-3PC) that is specifically designed to overcome this problem. In comparison with related models (e.g., LogGP) P-3PC is similar in complexity, but more accurate in many situations. Although the model is aimed at MPI's standard point-to-point operations, it is applicable to similar message passing definitions as well. The effectiveness of the model is tested in a framework for automatic parallelization of imaging applications. Experiments are performed on two Beowulf-type systems, each having a different interconnection network, and a different MPI implementation. Results show that, where other models frequently fail, P-3PC correctly predicts the communication costs related to any type of domain decomposition.

Dennis Koelma | Frank J. Seinstra | D. Koelma | F. Seinstra

[1] Francisco Tirado,et al. Data Locality Exploitation in the Decomposition of Regular Domain Problems , 2000, IEEE Trans. Parallel Distributed Syst..

[2] Csaba Andras Moritz,et al. LoGPC: Modeling Network Contention in Message-Passing Programs , 2001, IEEE Trans. Parallel Distributed Syst..

[3] Dennis Koelma,et al. A Software Architecture for User Transparent Parallel Image Processing on MIMD Computers , 2001, Euro-Par.

[4] Ramesh Subramonian,et al. LogP: towards a realistic model of parallel computation , 1993, PPOPP '93.

[5] Chris J. Scheiman,et al. LogGP: incorporating long messages into the LogP model—one step closer towards a realistic model for parallel computation , 1995, SPAA '95.

[6] F. J. Seinstra,et al. Modeling Performance of Low Level Image Processing Routines on MIMD Computers , 1999 .

[7] Henri E. Bal,et al. LFC: A Communication Substrate for Myrinet , 1998 .

[8] Amotz Bar-Noy,et al. Designing broadcasting algorithms in the postal model for message-passing systems , 2005, Mathematical systems theory.

[9] Peter M. A. Sloot,et al. The distributed ASCI Supercomputer project , 2000, OPSR.