نوع مقاله : مقاله پژوهشی
عنوان مقاله English
نویسندگان English
Automatic accent recognition in speech processing systems remains a challenging task in information-security applications, audio content analysis, and intelligent speech-based technologies. Among Arabic dialects, the Iraqi dialect holds particular importance due to its acoustic complexity, internal dialectal variation, and geopolitical significance. Nevertheless, many existing approaches, especially under noisy conditions, suffer from limited detection accuracy. In this paper, we propose an optimized method based on the Whisper model for detecting the Iraqi dialect among 17 Arabic dialects. In the proposed framework, speech signals from the ADI17 dataset are processed under operational conditions with an additive noise level corresponding to an SNR of 10 dB, simulating realistic acoustic environments. High-level acoustic features are extracted using the optimized Whisper model based on log-mel spectrogram representations, and these features are subsequently fed into a multi-layer perceptron (MLP) network for classification. To evaluate system performance, the dataset is divided into training, validation, and test subsets, and the problem is formulated as a multi-class classification task. Experimental results show that the proposed method achieves an innovative detection accuracy of 89% under 10 dB SNR conditions for Iraqi accent identification. Furthermore, the analysis of the confusion matrix and ROC curve demonstrates appropriate discriminative capability and stable performance in recognizing the target accent. The findings confirm that leveraging deep acoustic representations extracted by Whisper provides an effective and reliable approach for Iraqi accent recognition in noisy environments.
کلیدواژهها English