MPI Detach - Asynchronous Local Completion
Abstract
When aiming for large scale parallel computing, waiting time due to network latency, synchronization, and load imbalance are the primary opponents of high parallel efficiency. A common approach to hide latency with computation is the use of non-blocking communication. In the presence of a consistent load imbalance, synchronization cost is just the visible symptom of the load imbalance. Tasking approaches as in OpenMP, TBB, OmpSs, or C++20 coroutines promise to expose a higher degree of concurrency, which can be distributed on available execution units and significantly increase load balance. Available MPI non-blocking functionality does not integrate seamlessly into such tasking parallelization. In this work, we present a slim extension of the MPI interface to allow seamless integration of non-blocking communication with available concepts of asynchronous execution in OpenMP and C++.
Authors 5
-
Joachim Protze Aachen
Affiliation as printed
RWTH Aachen University ITC, Germany
-
Marc-André Hermanns Aachen
Affiliation as printed
RWTH Aachen University ITC, Germany
-
Ali Can Demiralp Aachen
Affiliation as printed
RWTH Aachen University VCI, Germany
-
Matthias Müller Aachen
Affiliation as printed
RWTH Aachen University ITC, Germany
-
Torsten Kuhlen Aachen
Affiliation as printed
RWTH Aachen University VCI, Germany
Cited by 9 stored of 9
9 results
No patents citing this paper on Lens.org (checked 2026-10-06).
References 9
-
W57462620details pending0citations
-
W1980670496details pending0citations
-
W1986190431details pending0citations
-
W2613247803details pending0citations
-
W2884677265details pending0citations
-
W2890814216details pending0citations
-
W3104104963details pending0citations
9 results