Adaptive Attention and Pyramid Pooling for High-Performance Dysarthric Speech Recognition
Speech pathology is a multidisciplinary field that assesses, diagnoses, and manages communication disorders caused by neurological, developmental, structural, and degenerative conditions. Among these disorders, dysarthria is a motor speech impairment characterized by reduced speech intelligibility and abnormal articulatory patterns resulting from neuromuscular dysfunction. Automatic dysarthria assessment remains challenging because of substantial inter-speaker variability, diverse severity levels, and the complex temporal–spectral characteristics of pathological speech signals. To address these challenges, this paper proposes an attention-enhanced convolutional neural network for automatic dysarthria detection and severity classification. The proposed architecture integrates a convolutional block attention module to enhance channel-wise and spatial feature representations, a coordinate attention module to capture long-range contextual dependencies while preserving positional information, and spatial pyramid pooling to extract robust multi-scale representations and accommodate input variability. Hierarchical convolutional layers with rectified linear unit activation and max-pooling operations progressively learn discriminative speech representations, followed by fully connected layers and a softmax classifier for prediction. The proposed framework was comprehensively evaluated on two publicly available datasets: TORGO Dataset 1, containing dysarthric and healthy control speech recordings, and TORGO Severity Dataset 2, comprising four dysarthria severity classes. Mel-frequency cepstral coefficients, together with their first- and second-order temporal derivatives and complementary acoustic descriptors, were extracted to characterize the speech signals. Extensive experiments, including cross-validation, repeated runs, ablation studies, Shapley additive explanations-based interpretability analysis, and statistical significance analysis, demonstrate the effectiveness and robustness of the proposed framework. On TORGO Dataset 1, the proposed model achieved 95.00% accuracy, 94.50% precision, 95.25% recall, and a 94.75% F1-score, outperforming conventional machine learning and deep learning baselines. On TORGO Severity Dataset 2, the proposed framework achieved 99.10% classification accuracy, an F1-score of 0.9910, and an area under the curve of 0.9999, demonstrating strong performance in multi-class dysarthria severity classification. These findings highlight the proposed model’s strong discriminative capability, robustness, and interpretability, indicating its potential as a reliable computer-aided decision-support system for dysarthria detection and severity assessment.

- Singh N, Tripathi P. Explainable SI-SAPN: signal imagery-based spatial attention pyramidal network model for Parkinson’s disease detection using speech spectrograms. Int J Inf Technol. 2025:1-21. doi: 10.1007/s41870-025-02693-9
- Zhang Z, Tang C, Yi W, Xu M, Occhipinti LG. Machine learning-driven smart socks with printed strain sensors for fall detection. In: 2025 IEEE BioSensors Conference (BioSensors). San Diego, CA, USA: IEEE; 2025:1-4. doi: 10.1109/BioSensors65002.2025.11239179
- Wen-Tsai S, Hao-Wei K, Sung-Jung H. Speech recognition via CTC-CNN model. Comput Mater Contin. 2023;76(3):3833-3848. doi: 10.32604/cmc.2023.040024
- Singh S, Wang Q, Zhong Z, et al. Robust Cross-Etiology and Speaker-Independent Dysarthric Speech Recognition. arXiv. Preprint posted online 2025. doi: 10.48550/arXiv.2501.14994
- Spencer KA, Eddy B, Papathanasiou I, Summers D, Britton D. Management of velopharyngeal impairment in adults with dysarthria: a systematic review. Am J Speech Lang Pathol. 2025;34(1):391-409. doi: 10.1044/2024_AJSLP-24-00287
- Hao J, Pan H. A novel deep transformer-based CvT model for sign language recognition in visual communication. Sci Rep. 2025:16(1). doi: 10.1038/s41598-025-31558-1
- Hasan MT, Hossain MAE, Mukta MSH, Akter A, Ahmed M, Islam S. A review on deep learning-based cyberbullying detection. Future Internet. 2023;15(5):179. doi: 10.3390/fi15050179
- Azad M, Khan MFK, El-Ghany SA. XAI-enhanced machine learning for obesity risk classification: a stacking approach with LIME explanations. IEEE Access. 2025;13:13847-13865. doi: 10.1109/ACCESS.2025.3530840
- Shams MY, Tarek Z, Elshewey AM. A novel RFE-GRU model for diabetes classification using PIMA Indian dataset. Sci Rep. 2025;15(1):982. doi: 10.1038/s41598-024-82420-9
- Liu S, Geng M, Hu S, et al. Recent progress in the CUHK dysarthric speech recognition system. IEEE/ACM Trans Audio Speech Lang Process. 2021;29:2267-2281. doi: 10.1109/TASLP.2021.3091805
- Hassan E, Elbedwehy S, Shams MY, Abd El-Hafeez T, El-Rashidy N. Optimizing poultry audio signal classification with deep learning and burn layer fusion. J Big Data. 2024;11(1):135. doi: 10.1186/s40537-024-00985-8
- Qian Z, Xiao K, Yu C. A survey of technologies for automatic dysarthric speech recognition. EURASIP J Audio Speech Music Process. 2023;2023(1):48. doi: 10.1186/s13636-023-00318-2
- He Y, Seng KP, Lim CS, Ang LM. Robust dysarthric speech recognition with GAN enhancement and LLM correction. Adv Intell Syst. 2026;8(2):e202500873. doi: 10.1002/aisy.202500873
- Hertrich I, Dietrich S, Blum C, Ackermann H. The role of the dorsolateral prefrontal cortex for speech and language processing. Front Hum Neurosci. 2021;15:645209, doi: 10.3389/fnhum.2021.645209
- Basilakos A, Fridriksson J. Types of motor speech impairments associated with neurologic diseases. In: Handbook of Clinical Neurology. Elsevier; 2022:71-79. doi: 10.1016/b978-0-12-823384-9.00004-9
- Chung Y, Hong J, Lee J, Kim E. Diagnosis-aware multitask fine-tuning of Whisper for dysarthric speech recognition. Speech Commun. 2026;180,103393. doi: 10.1016/j.specom.2026.103393
- Choi J, Moya-Galé G, Hwang K, Hirschberg J, Levy ES. Automatic speech recognition for intelligibility assessment in children with dysarthria. J Speech Lang Hear Res. 2026;69(4):1438-1454. doi: 10.1044/2025_JSLHR-25-00562
- Kim Y, Kent RD, Weismer G. An acoustic study of the relationships among neurologic disease, dysarthria type, and severity of dysarthria. J Speech Lang Hear Res. 2011;54(2):417-429. doi: 10.1044/1092-4388(2010/10-0020)
- Narendra NP, Alku P. Glottal source information for pathological voice detection. IEEE Access. 2020;8:67745-67755. doi: 10.1109/ACCESS.2020.2986171
- Joshy AA, Rajan R. Automated dysarthria severity classification: a study on acoustic features and deep learning techniques. IEEE Trans Neural Syst Rehabil Eng. 2022;30:1147-1157. doi: 10.1109/TNSRE.2022.3169814
- Javanmardi F, Kadiri SR, Alku P. Pre-trained models for detection and severity level classification of dysarthria from speech. Speech Commun. 2024;158:103047. doi: 10.1016/j.specom.2024.103047
- Stumpf L, Kadirvelu B, Waibel S, Faisal AA. Speaker-independent dysarthria severity classification using self-supervised transformers and multi-task learning. arXiv. 2024. doi: 10.48550/arXiv.2403.00854
- Shabber SM, Sumesh EP. AFM signal model for dysarthric speech classification using speech biomarkers. Front Hum Neurosci. 2024;18:1346297. doi: 10.3389/fnhum.2024.1346297
- Radha K, Bansal M. Automated detection and severity assessment of dysarthria using raw speech. In: 2023 14th International Conference on Computing, Communication and Networking Technologies (ICCCNT). Delhi, India: IEEE; 2023:1-7. doi: 10.1109/ICCCNT56998.2023.10307923
- Radha K, Bansal M, Dulipalla VR. Variable STFT layered CNN model for automated dysarthria detection and severity assessment using raw speech. Circuits Syst Signal Process. 2024;43(5):3261-3278. doi: 10.1007/s00034-024-02611-7
- Shanmugapriya P, Mohan V. Comparative analysis of deep learning models for dysarthric speech detection. Soft Comput. 2023;28(6):5683-5698. doi: 10.1007/s00500-023-09302-6
- Sekhar SM, Kashyap G, Bhansali A, Singh K. Dysarthric speech detection using transfer learning with convolutional neural networks. ICT Express. 2022;8(1):61-64. doi: 10.1016/j.icte.2021.07.004
- Alharbi GF, Alamri NK, Sabbeh SF. Automatic classification of speech dysarthric intelligibility levels using textual feature. IEEE Access. 2025;13:39982-39992. doi: 10.1109/ACCESS.2025.3547041
- S A, R P, M R. Comparative analysis of different time-frequency image representations for the detection and severity classification of dysarthric speech using deep learning. Results Eng. 2025;25:104561. doi: 10.1016/j.rineng.2025.104561
- Al-Ali AS, Haris RM, Akbari Y, Saleh M, Al-Maadeed S, Kumar MR. Integrating binary classification and clustering for multi-class dysarthria severity level classification: a two-stage approach. Cluster Comput. 2025;28(2):136. doi: 10.1007/s10586-024-04748-1
- Shabber SM, Sumesh EP, Ramachandran VL. Scalogram-based performance comparison of deep learning architectures for dysarthric speech detection. Artif Intell Rev. 2025;58(5):128. doi: 10.1007/s10462-024-11085-7
- Wang Q, Zhong Z, Singh S, et al. Dysarthric Speech Conformer: adaptation for sequence-to-sequence dysarthric speech recognition. In: ICASSP 2025–2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Hyderabad, India: IEEE; 2025:1-5. doi: 10.1109/ICASSP49660.2025.10889046
- Enderby P. Disorders of communication. In: Handbook of Clinical Neurology. Amsterdam, The Netherlands: Elsevier; 2013:273-281. doi: 10.1016/B978-0-444-52901-5.00022-8
- Zhang Z, Wang M. Convolutional neural network with convolutional block attention module for finger vein recognition. arXiv. Preprint posted online 2022. doi: 10.48550/arXiv.2202.06673
- Zakariah M, Alnuaim A. Recognizing human activities with the use of convolutional block attention module. Egypt Inform J. 2024;27:100536. doi: 10.1016/j.eij.2024.100536
- Aryal N, Lee SW. Frequency-based CNN and attention module for acoustic scene classification. Appl Acoust. 2023;210:109411. doi: 10.1016/j.apacoust.2023.109411
- Woo S, Park J, Lee JY, Kweon IS. CBAM: Convolutional Block Attention Module. In: Computer Vision – ECCV 2018 (Lecture Notes in Computer Science). Cham, Switzerland: Springer International Publishing; 2018:3-19. doi: 10.1007/978-3-030-01234-2_1
- Chen L, Yao H, Fu J, Ng CT. The classification and localization of crack using lightweight convolutional neural network with CBAM. Eng Struct. 2023;275:115291. doi: 10.1016/j.engstruct.2022.115291
- Rokhva S, Teimourpour B. Accurate and real-time food classification through the synergistic integration of EfficientNetB7, CBAM, transfer learning, and data augmentation. Food Humanit. 2025;4:100492. doi: 10.1016/j.foohum.2024.100492
- Hassan E, Shams MY, Hikal NA, Elmougy S. A novel convolutional neural network model for malaria cell images classification. Comput Mater Contin. 2022;72(3):5889-5907. doi: 10.32604/cmc.2022.025629
- Sarhan S, Nasr AA, Shams MY. Multipose Face Recognition-Based Combined Adaptive Deep Learning Vector Quantization. Comput Intell Neurosci. 2020;2020:1-11. doi: 10.1155/2020/8821868
- Hassan E, Shams MY, Hikal NA, Elmougy S. The effect of choosing optimizer algorithms to improve computer vision tasks: a comparative study. Multimed Tools Appl. 2023;82(11):16591-16633. doi: 10.1007/s11042-022-13820-0
- Joshy AA, Parameswaran PN, Janardan S, TK T, Rajan R. How dysarthria severity influences speech recognition in cerebral palsy. J Neurolinguistics. 2026;79:101348, doi: 10.1016/j.jneuroling.2026.101348
- Altulyan M. A robust and integrated speech recognition tool for dysarthria patients using lip movement recognition. Eng Technol Appl Sci Res. 2026;16(3):35660-35669. doi: 10.48084/etasr.17799
