论文信息 - Parallel Reproducible Summation

Parallel Reproducible Summation

Reproducibility, i.e. getting bitwise identical floating point results from multiple runs of the same program, is a property that many users depend on either for debugging or correctness checking in many codes [10]. However, the combination of dynamic scheduling of parallel computing resources, and floating point nonassociativity, makes attaining reproducibility a challenge even for simple reduction operations like computing the sum of a vector of numbers in parallel. We propose a technique for floating point summation that is reproducible independent of the order of summation. Our technique uses Rump's algorithm for error-free vector transformation [7], and is much more efficient than using (possibly very) high precision arithmetic. Our algorithm reproducibly computes highly accurate results with an absolute error bound of n · 2-28 macheps maxiIviI at a cost of 7n FLOPs and a small constant amount of extra memory usage. Higher accuracies are also possible by increasing the number of error-free transformations. As long as all operations are performed in to-nearest rounding mode, results computed by the proposed algorithms are reproducible for any run on any platform. In particular, our algorithm requires the minimum number of reductions, i.e. one reduction of an array of six double precision floating point numbers per sum, and hence is well suited for massively parallel environments.

James Demmel | Hong Diep Nguyen | J. Demmel

[1] Siegfried M. Rump,et al. Ultimately Fast Accurate Summation , 2009, SIAM J. Sci. Comput..

[2] Sriram Krishnamoorthy,et al. Effects of floating-point non-associativity on numerical computations on massively multithreaded systems , 2009 .

[3] Yozo Hida,et al. Accurate Floating Point Summation , 2002 .

[4] T. J. Dekker,et al. A floating-point technique for extending the available precision , 1971 .

[5] Jürgen Wolff von Gudenberg,et al. A long accumulator like a carry-save adder , 2011, Computing.

[6] Siegfried M. Rump,et al. Fast high precision summation , 2010 .

[7] James Demmel,et al. Fast Reproducible Floating-Point Summation , 2013, 2013 IEEE 21st Symposium on Computer Arithmetic.

[8] Philip Saponaro,et al. Improving numerical reproducibility and stability in large-scale numerical simulations on GPUs , 2010, 2010 IEEE International Symposium on Parallel & Distributed Processing (IPDPS).

[9] Nicholas J. Higham,et al. INVERSE PROBLEMS NEWSLETTER , 1991 .

[10] Wayne B. Hayes,et al. Algorithm 908 , 2010 .