A

GeHirNet: A Gender-Aware Hierarchical Model for Voice Pathology Classification

arXiv (Cornell University)

Abstract

AI-based voice analysis shows promise for disease diagnostics, but existing classifiers often fail to accurately identify specific pathologies because of gender-related acoustic variations and the scarcity of data for rare diseases. We propose a novel two-stage framework that first identifies gender-specific pathological patterns using ResNet-50 on Mel spectrograms, then performs gender-conditioned disease classification. We address class imbalance through multi-scale resampling and time warping augmentation. Evaluated on a merged dataset from four public repositories, our two-stage architecture with time warping achieves state-of-the-art performance (97.63\% accuracy, 95.25\% MCC), with a 5\% MCC improvement over single-stage baseline. This work advances voice pathology classification while reducing gender bias through hierarchical modeling of vocal characteristics.

Authors 4

  1. ETH Zurich

    Affiliation as printed

    Centre for Digital Health Interventions, ETH Zurich, Zurich, Switzerland

  2. RWTH Aachen University

    Affiliation as printed

    Institute of Mechanism Theory, Machine Dynamics and Robotics, RWTH Aachen University, Aachen, Germany

  3. ETH Zurich · University of St.Gallen

    Affiliation as printed

    Centre for Digital Health Interventions, ETH Zurich, Zurich, Switzerland

    Centre for Digital Health Interventions, University of St. Gallen, St. Gallen, Switzerland

  4. ETH Zurich

    Affiliation as printed

    Centre for Digital Health Interventions, ETH Zurich, Zurich, Switzerland

Cited by 1 stored of 1

1 result

No patents citing this paper on Lens.org (checked 2026-10-06).

References 0