A

Configuring Large Reasoning Models Using Process Mining: A Benchmark and a Case Study

Lecture notes in business information processing, pp. 412–424

Abstract

Large Reasoning Models (LRMs), a subset of Large Language Models (LLMs) trained to articulate their chain-of-thought, have shown promise in tackling complex scientific tasks. However, evaluating and configuring their reasoning processes remains underexplored. This paper leverages a process mining-specific LLM evaluation framework to propose a methodology for analyzing and configuring LRMs. We introduce an approach to extract and classify reasoning steps by type (e.g., Deductive Reasoning, or Hypothesis Generation) and effect (Positive, Indifferent, Negative) on the overall reasoning, enabling a detailed assessment of reasoning quality. From this, we derive a new benchmark, PMLRM-Bench, which evaluates not only the correctness of outputs but also the robustness of the reasoning process. A case study on the QwQ-32B LLM demonstrates how targeted adjustments to reasoning type frequencies can boost task-specific performance. Our results reveal distinct reasoning patterns across models and provide actionable insights for LRM configuration. This work bridges process mining and LLM evaluation, offering a scalable framework for reasoning analysis.

Authors 4

  1. RWTH Aachen University

    Affiliation as printed

    Process and Data Science Chair, RWTH Aachen University

  2. RWTH Aachen University

    Affiliation as printed

    Process and Data Science Chair, RWTH Aachen University

  3. RWTH Aachen University

    Affiliation as printed

    Process and Data Science Chair, RWTH Aachen University

  4. RWTH Aachen University

    Affiliation as printed

    Process and Data Science Chair, RWTH Aachen University

Cited by 2 stored of 2

2 results

No patents citing this paper on Lens.org (checked 2026-10-06).

References 0