Designing Quality MPI Correctness Benchmarks: Insights and Metrics
Abstract
Several MPI correctness benchmarks have been proposed to evaluate the quality of MPI correctness tools. The design of such a benchmark comes with different challenges, which we address in this paper. First, an imbalance in the proportion of correct and erroneous codes in the benchmarks requires careful metric interpretation (recall, accuracy, F1 score). Second, tools that detect errors but do not report additional information, like the affected source line or class of error, are less valuable. We extend the typical notion of a true positive with stricter variants that consider a tool’s helpfulness. We introduce a new noise metric to consider the amount of distracting error reports. We evaluate those new metrics with MPI-BugBench, on the MPI correctness tools ITAC, MUST, and PARCOACH. Third, we discuss the complexities of hand-crafted and automatically generated benchmark codes and the additional challenges of non-deterministic errors.
Authors 7
-
Technische Universität Darmstadt
Affiliation as printed
Technical University Darmstadt
-
Simon Schwitanski Aachen
Affiliation as printed
RWTH Aachen University
-
Institut national de recherche en sciences et technologies du numérique
Affiliation as printed
Inria
-
Technische Universität Darmstadt
Affiliation as printed
Technical University Darmstadt
-
Joachim Jenke Aachen
Affiliation as printed
RWTH Aachen University
-
Affiliation as printed
Eviden
-
Technische Universität Darmstadt
Affiliation as printed
Technical University Darmstadt
Cited by 1 stored of 1
1 result
No patents citing this paper on Lens.org (checked 2026-10-06).
References 15
-
W1647645866details pending0citations
-
W2985025837details pending0citations
-
W4200404442details pending0citations
-
W4200631433details pending0citations
-
W1565978235details pending0citations
-
W2479360701details pending0citations
-
W2969137362details pending0citations
15 results