A

Applying Process Mining on Scientific Workflows: a Case Study on High Performance Computing Data

arXiv (Cornell University)

Abstract

Computer-based scientific experiments are becoming increasingly data-intensive, necessitating the use of High-Performance Computing (HPC) clusters to handle large scientific workflows. These workflows result in complex data and control flows within the system, making analysis challenging. This paper focuses on the extraction of case IDs from SLURM-based HPC cluster logs, a crucial step for applying mainstream process mining techniques. The core contribution is the development of methods to correlate jobs in the system, whether their interdependencies are explicitly specified or not. We present our log extraction and correlation techniques, supported by experiments that validate our approach, enabling comprehensive documentation of workflows and identification of performance bottlenecks.

Authors 4

  1. RWTH Aachen University

    Affiliation as printed

    Chair of Process and Data Science , RWTH Aachen University , Aachen , Germany

  2. RWTH Aachen University

    Affiliation as printed

    Chair of Process and Data Science , RWTH Aachen University , Aachen , Germany

  3. RWTH Aachen University

    Affiliation as printed

    Chair of Process and Data Science , RWTH Aachen University , Aachen , Germany

  4. RWTH Aachen University

    Affiliation as printed

    Chair of Process and Data Science , RWTH Aachen University , Aachen , Germany

Cited by 1 stored of 1

1 result

No patents citing this paper on Lens.org (checked 2026-10-06).

References 0