A

Designing Quality MPI Correctness Benchmarks: Insights and Metrics

Abstract

Several MPI correctness benchmarks have been proposed to evaluate the quality of MPI correctness tools. The design of such a benchmark comes with different challenges, which we address in this paper. First, an imbalance in the proportion of correct and erroneous codes in the benchmarks requires careful metric interpretation (recall, accuracy, F1 score). Second, tools that detect errors but do not report additional information, like the affected source line or class of error, are less valuable. We extend the typical notion of a true positive with stricter variants that consider a tool’s helpfulness. We introduce a new noise metric to consider the amount of distracting error reports. We evaluate those new metrics with MPI-BugBench, on the MPI correctness tools ITAC, MUST, and PARCOACH. Third, we discuss the complexities of hand-crafted and automatically generated benchmark codes and the additional challenges of non-deterministic errors.

Authors 7

  1. Technische Universität Darmstadt

    Affiliation as printed

    Technical University Darmstadt

  2. RWTH Aachen University

    Affiliation as printed

    RWTH Aachen University

  3. Institut national de recherche en sciences et technologies du numérique

    Affiliation as printed

    Inria

  4. Technische Universität Darmstadt

    Affiliation as printed

    Technical University Darmstadt

  5. Joachim Jenke Aachen

    RWTH Aachen University

    Affiliation as printed

    RWTH Aachen University

  6. Affiliation as printed

    Eviden

  7. Technische Universität Darmstadt

    Affiliation as printed

    Technical University Darmstadt

Cited by 1 stored of 1

1 result

No patents citing this paper on Lens.org (checked 2026-10-06).

References 15

15 results