Gender Detection Using Short Speech Utterances: A Deep Learning Approach with Wavelet Cepstral Coefficients

Closed

Syahroni Hidayat, Budi Sunarko, Uswatun Hasanah, Faila Nadhifatul Aryza

2024 7th International Seminar on Research of Information Technology and Intelligent Systems: Advanced Intelligent Systems in Contemporary Society, ISRITI 2024 - Proceedings Conference paper Cited by 1 Quartile

Abstract

The human voice is critical in recognizing the speaker's identity and gender. This study aims to develop an effective gender detection system using short utterances of less than one second. Wavelet Cepstral Coefficient (WCC) features were used for feature extraction, and a Long Short-Term Memory (LSTM) model was applied for classification. Primary and secondary short voice datasets, including syllables and the Audio MNIST dataset, underwent acoustic pretreatment and were trained using LSTM architectures with 32, 64, and 128 units. Results indicated that the LSTM model achieved high accuracy, with the syllable dataset reaching 0.99 accuracy and the Audio MNIST dataset showing perfect accuracy (1.00) using a 32-unit configuration. The findings demonstrate that combining WCC features and LSTM models can efficiently handle short-duration speech for accurate gender detection. © 2024 IEEE.

Affiliations

Universitas Negeri Semarang, Dept. of Electrical Engineering, Semarang, Indonesia; Universitas Negeri Semarang, Dept. of Informatics and Computer Engineering Education, Semarang, Indonesia; Universitas Negeri Semarang, Dept. of Computer Engineering, Semarang, Indonesia