AccScience Publishing / AIH / Online First / DOI: 10.36922/AIH026280086
Cite this article
23
Views
Related Info Links
More by Authors Links
Journal Browser
Volume | Year
Issue
Search
News and Announcements
View All
ORIGINAL RESEARCH ARTICLE

Adaptive Attention and Pyramid Pooling for High-Performance Dysarthric Speech Recognition

Zahraa Tarek1 Tarek Abd El-Hafeez El-Hafeez2 Esraa Hassan3* Ahmed Fahim4 Khaled Alnowaiser4 Mahmoud Y. Shams3
Show Less
1 Department of Computer Engineering and Information, College of Engineering, Wadi Ad Dwaser, Prince Sattam Bin Abdulaziz University, Al-Kharj 16273, Saudi Arabia, Al-Kharj , Saudi Arabia
2 Department of Computer Science, Faculty of Science, Minia University, Minia , Egypt
3 Department of Machine Learning and Information Retrieval, Faculty of Artificial Intelligence, Kafrelsheikh University, Kafrelsheikh , Egypt
4 Department of Computer Science, College of Computer Engineering and Sciences, Prince Sattam Bin Abdulaziz University , Al-Kharj
Received: 10 July 2026 | Revised: 27 July 2026 | Accepted: 6 August 2026 | Published online: 4 September 2026
© 2026 by the Author(s). This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution 4.0 International License ( https://creativecommons.org/licenses/by/4.0/ )
Abstract

Speech pathology is a multidisciplinary field that assesses, diagnoses, and manages communication disorders caused by neurological, developmental, structural, and degenerative conditions. Among these disorders, dysarthria is a motor speech impairment characterized by reduced speech intelligibility and abnormal articulatory patterns resulting from neuromuscular dysfunction. Automatic dysarthria assessment remains challenging because of substantial inter-speaker variability, diverse severity levels, and the complex temporal–spectral characteristics of pathological speech signals. To address these challenges, this paper proposes an attention-enhanced convolutional neural network for automatic dysarthria detection and severity classification. The proposed architecture integrates a convolutional block attention module to enhance channel-wise and spatial feature representations, a coordinate attention module to capture long-range contextual dependencies while preserving positional information, and spatial pyramid pooling to extract robust multi-scale representations and accommodate input variability. Hierarchical convolutional layers with rectified linear unit activation and max-pooling operations progressively learn discriminative speech representations, followed by fully connected layers and a softmax classifier for prediction. The proposed framework was comprehensively evaluated on two publicly available datasets: TORGO Dataset 1, containing dysarthric and healthy control speech recordings, and TORGO Severity Dataset 2, comprising four dysarthria severity classes. Mel-frequency cepstral coefficients, together with their first- and second-order temporal derivatives and complementary acoustic descriptors, were extracted to characterize the speech signals. Extensive experiments, including cross-validation, repeated runs, ablation studies, Shapley additive explanations-based interpretability analysis, and statistical significance analysis, demonstrate the effectiveness and robustness of the proposed framework. On TORGO Dataset 1, the proposed model achieved 95.00% accuracy, 94.50% precision, 95.25% recall, and a 94.75% F1-score, outperforming conventional machine learning and deep learning baselines. On TORGO Severity Dataset 2, the proposed framework achieved 99.10% classification accuracy, an F1-score of 0.9910, and an area under the curve of 0.9999, demonstrating strong performance in multi-class dysarthria severity classification. These findings highlight the proposed model’s strong discriminative capability, robustness, and interpretability, indicating its potential as a reliable computer-aided decision-support system for dysarthria detection and severity assessment.

Graphical abstract
Keywords
Dysarthria detection
TORGO database
Convolutional neural networks
Attention mechanisms
Speech pathology
Machine learning
Funding
The authors extend their appreciation to Prince Sattam bin Abdulaziz University for funding this research work through the project number PSAU/2025/01/34420.
Conflict of interest
The authors declare no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
References
  1. Singh N, Tripathi P. Explainable SI-SAPN: signal imagery-based spatial attention pyramidal network model for Parkinson’s disease detection using speech spectrograms. Int J Inf Technol. 2025:1-21. doi: 10.1007/s41870-025-02693-9
  2. Zhang Z, Tang C, Yi W, Xu M, Occhipinti LG. Machine learning-driven smart socks with printed strain sensors for fall detection. In: 2025 IEEE BioSensors Conference (BioSensors). San Diego, CA, USA: IEEE; 2025:1-4. doi: 10.1109/BioSensors65002.2025.11239179
  3. Wen-Tsai S, Hao-Wei K, Sung-Jung H. Speech recognition via CTC-CNN model. Comput Mater Contin. 2023;76(3):3833-3848. doi: 10.32604/cmc.2023.040024
  4. Singh S, Wang Q, Zhong Z, et al. Robust Cross-Etiology and Speaker-Independent Dysarthric Speech Recognition. arXiv. Preprint posted online 2025. doi: 10.48550/arXiv.2501.14994
  5. Spencer KA, Eddy B, Papathanasiou I, Summers D, Britton D. Management of velopharyngeal impairment in adults with dysarthria: a systematic review. Am J Speech Lang Pathol. 2025;34(1):391-409. doi: 10.1044/2024_AJSLP-24-00287
  6. Hao J, Pan H. A novel deep transformer-based CvT model for sign language recognition in visual communication. Sci Rep. 2025:16(1). doi: 10.1038/s41598-025-31558-1
  7. Hasan MT, Hossain MAE, Mukta MSH, Akter A, Ahmed M, Islam S. A review on deep learning-based cyberbullying detection. Future Internet. 2023;15(5):179. doi: 10.3390/fi15050179
  8. Azad M, Khan MFK, El-Ghany SA. XAI-enhanced machine learning for obesity risk classification: a stacking approach with LIME explanations. IEEE Access. 2025;13:13847-13865. doi: 10.1109/ACCESS.2025.3530840
  9. Shams MY, Tarek Z, Elshewey AM. A novel RFE-GRU model for diabetes classification using PIMA Indian dataset. Sci Rep. 2025;15(1):982. doi: 10.1038/s41598-024-82420-9
  10. Liu S, Geng M, Hu S, et al. Recent progress in the CUHK dysarthric speech recognition system. IEEE/ACM Trans Audio Speech Lang Process. 2021;29:2267-2281. doi: 10.1109/TASLP.2021.3091805
  11. Hassan E, Elbedwehy S, Shams MY, Abd El-Hafeez T, El-Rashidy N. Optimizing poultry audio signal classification with deep learning and burn layer fusion. J Big Data. 2024;11(1):135. doi: 10.1186/s40537-024-00985-8
  12. Qian Z, Xiao K, Yu C. A survey of technologies for automatic dysarthric speech recognition. EURASIP J Audio Speech Music Process. 2023;2023(1):48. doi: 10.1186/s13636-023-00318-2
  13. He Y, Seng KP, Lim CS, Ang LM. Robust dysarthric speech recognition with GAN enhancement and LLM correction. Adv Intell Syst. 2026;8(2):e202500873. doi: 10.1002/aisy.202500873
  14. Hertrich I, Dietrich S, Blum C, Ackermann H. The role of the dorsolateral prefrontal cortex for speech and language processing. Front Hum Neurosci. 2021;15:645209, doi: 10.3389/fnhum.2021.645209
  15. Basilakos A, Fridriksson J. Types of motor speech impairments associated with neurologic diseases. In: Handbook of Clinical Neurology. Elsevier; 2022:71-79. doi: 10.1016/b978-0-12-823384-9.00004-9
  16. Chung Y, Hong J, Lee J, Kim E. Diagnosis-aware multitask fine-tuning of Whisper for dysarthric speech recognition. Speech Commun. 2026;180,103393. doi: 10.1016/j.specom.2026.103393
  17. Choi J, Moya-Galé G, Hwang K, Hirschberg J, Levy ES. Automatic speech recognition for intelligibility assessment in children with dysarthria. J Speech Lang Hear Res. 2026;69(4):1438-1454. doi: 10.1044/2025_JSLHR-25-00562
  18. Kim Y, Kent RD, Weismer G. An acoustic study of the relationships among neurologic disease, dysarthria type, and severity of dysarthria. J Speech Lang Hear Res. 2011;54(2):417-429. doi: 10.1044/1092-4388(2010/10-0020)
  19. Narendra NP, Alku P. Glottal source information for pathological voice detection. IEEE Access. 2020;8:67745-67755. doi: 10.1109/ACCESS.2020.2986171
  20. Joshy AA, Rajan R. Automated dysarthria severity classification: a study on acoustic features and deep learning techniques. IEEE Trans Neural Syst Rehabil Eng. 2022;30:1147-1157. doi: 10.1109/TNSRE.2022.3169814
  21. Javanmardi F, Kadiri SR, Alku P. Pre-trained models for detection and severity level classification of dysarthria from speech. Speech Commun. 2024;158:103047. doi: 10.1016/j.specom.2024.103047
  22. Stumpf L, Kadirvelu B, Waibel S, Faisal AA. Speaker-independent dysarthria severity classification using self-supervised transformers and multi-task learning. arXiv. 2024. doi: 10.48550/arXiv.2403.00854
  23. Shabber SM, Sumesh EP. AFM signal model for dysarthric speech classification using speech biomarkers. Front Hum Neurosci. 2024;18:1346297. doi: 10.3389/fnhum.2024.1346297
  24. Radha K, Bansal M. Automated detection and severity assessment of dysarthria using raw speech. In: 2023 14th International Conference on Computing, Communication and Networking Technologies (ICCCNT). Delhi, India: IEEE; 2023:1-7. doi: 10.1109/ICCCNT56998.2023.10307923
  25. Radha K, Bansal M, Dulipalla VR. Variable STFT layered CNN model for automated dysarthria detection and severity assessment using raw speech. Circuits Syst Signal Process. 2024;43(5):3261-3278. doi: 10.1007/s00034-024-02611-7
  26. Shanmugapriya P, Mohan V. Comparative analysis of deep learning models for dysarthric speech detection. Soft Comput. 2023;28(6):5683-5698. doi: 10.1007/s00500-023-09302-6
  27. Sekhar SM, Kashyap G, Bhansali A, Singh K. Dysarthric speech detection using transfer learning with convolutional neural networks. ICT Express. 2022;8(1):61-64. doi: 10.1016/j.icte.2021.07.004
  28. Alharbi GF, Alamri NK, Sabbeh SF. Automatic classification of speech dysarthric intelligibility levels using textual feature. IEEE Access. 2025;13:39982-39992. doi: 10.1109/ACCESS.2025.3547041
  29. S A, R P, M R. Comparative analysis of different time-frequency image representations for the detection and severity classification of dysarthric speech using deep learning. Results Eng. 2025;25:104561. doi: 10.1016/j.rineng.2025.104561
  30. Al-Ali AS, Haris RM, Akbari Y, Saleh M, Al-Maadeed S, Kumar MR. Integrating binary classification and clustering for multi-class dysarthria severity level classification: a two-stage approach. Cluster Comput. 2025;28(2):136. doi: 10.1007/s10586-024-04748-1
  31. Shabber SM, Sumesh EP, Ramachandran VL. Scalogram-based performance comparison of deep learning architectures for dysarthric speech detection. Artif Intell Rev. 2025;58(5):128. doi: 10.1007/s10462-024-11085-7
  32. Wang Q, Zhong Z, Singh S, et al. Dysarthric Speech Conformer: adaptation for sequence-to-sequence dysarthric speech recognition. In: ICASSP 2025–2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Hyderabad, India: IEEE; 2025:1-5. doi: 10.1109/ICASSP49660.2025.10889046
  33. Enderby P. Disorders of communication. In: Handbook of Clinical Neurology. Amsterdam, The Netherlands: Elsevier; 2013:273-281. doi: 10.1016/B978-0-444-52901-5.00022-8
  34. Zhang Z, Wang M. Convolutional neural network with convolutional block attention module for finger vein recognition. arXiv. Preprint posted online 2022. doi: 10.48550/arXiv.2202.06673
  35. Zakariah M, Alnuaim A. Recognizing human activities with the use of convolutional block attention module. Egypt Inform J. 2024;27:100536. doi: 10.1016/j.eij.2024.100536
  36. Aryal N, Lee SW. Frequency-based CNN and attention module for acoustic scene classification. Appl Acoust. 2023;210:109411. doi: 10.1016/j.apacoust.2023.109411
  37. Woo S, Park J, Lee JY, Kweon IS. CBAM: Convolutional Block Attention Module. In: Computer Vision – ECCV 2018 (Lecture Notes in Computer Science). Cham, Switzerland: Springer International Publishing; 2018:3-19. doi: 10.1007/978-3-030-01234-2_1
  38. Chen L, Yao H, Fu J, Ng CT. The classification and localization of crack using lightweight convolutional neural network with CBAM. Eng Struct. 2023;275:115291. doi: 10.1016/j.engstruct.2022.115291
  39. Rokhva S, Teimourpour B. Accurate and real-time food classification through the synergistic integration of EfficientNetB7, CBAM, transfer learning, and data augmentation. Food Humanit. 2025;4:100492. doi: 10.1016/j.foohum.2024.100492
  40. Hassan E, Shams MY, Hikal NA, Elmougy S. A novel convolutional neural network model for malaria cell images classification. Comput Mater Contin. 2022;72(3):5889-5907. doi: 10.32604/cmc.2022.025629
  41. Sarhan S, Nasr AA, Shams MY. Multipose Face Recognition-Based Combined Adaptive Deep Learning Vector Quantization. Comput Intell Neurosci. 2020;2020:1-11. doi: 10.1155/2020/8821868
  42. Hassan E, Shams MY, Hikal NA, Elmougy S. The effect of choosing optimizer algorithms to improve computer vision tasks: a comparative study. Multimed Tools Appl. 2023;82(11):16591-16633. doi: 10.1007/s11042-022-13820-0
  43. Joshy AA, Parameswaran PN, Janardan S, TK T, Rajan R. How dysarthria severity influences speech recognition in cerebral palsy. J Neurolinguistics. 2026;79:101348, doi: 10.1016/j.jneuroling.2026.101348
  44. Altulyan M. A robust and integrated speech recognition tool for dysarthria patients using lip movement recognition. Eng Technol Appl Sci Res. 2026;16(3):35660-35669. doi: 10.48084/etasr.17799
Share
Back to top
Artificial Intelligence in Health, Electronic ISSN: 3029-2387 Print ISSN: 3041-0894, Published by AccScience Publishing