Celebrating 30 Years

Research Progress and Trends of Intelligent Speech in Pathological Healthcare

Expand
  • 1. School of Artificial Intelligence, Beijing University of Posts and Telecommunications, Beijing 100876, China; 2. School of Ecoled’Ing′enieurs Paris, Shanghai Jiao Tong University, Shanghai 200240, China; 3. X-LANCE Lab, MoE Key Lab of Artificial Intelligence, School of Computer Science, Shanghai Jiao Tong University, Shanghai 200240, China

Received date: 2026-01-21

  Revised date: 2026-05-11

  Accepted date: 2026-05-21

  Online published: 2026-06-20

Abstract

Intelligent speech technology, rooted in the ancient medical practice of “auscultation and interrogation”, is emerging as a transformative tool in modern pathological healthcare. By analyzing acoustic biomarkers within speech, voice, cough, breath sounds, and heart sounds, it offers a non-invasive, cost-effective avenue for early screening, auxiliary diagnosis, monitoring, and rehabilitation assessment across a wide spectrum of conditions, including mental disorders, neurodegenerative diseases, respiratory illnesses, cardiovascular diseases, and laryngeal or vocal tract pathologies. This study comprehensively reviews the research progress and prevailing trends in this interdisciplinary field. It begins by elucidating the physiological basis of pathological acoustics, including neural dysregulation, airway and pulmonary abnormalities, hemodynamic disturbances, and laryngeal or vocal tract dysfunction, and then discusses its integration with AI-driven diagnostics. The core of the review systematically details advances in two pillars: data resource construction (encompassing datasets for various diseases and standardization efforts) and methodological innovation (tracking the paradigm shift from feature-based machine learning to deep learning, self-supervised models, and multimodal large language models). Furthermore, it explores the development of intelligent speech-driven intervention systems for mental health. The analysis identifies key dynamic trends: the evolution from single-modality to multimodal analysis, the shift from strong to weak/self-supervised learning, the transition from controlled lab settings to naturalistic scenarios, the growing priority of model interpretability, and the move towards multi-disease coexistence modeling. Despite promising clinical potential, significant challenges persist, including data scarcity, algorithmic robustness, and clinical integration bottlenecks. The study concludes by outlining critical future directions: fostering federated learning and multi-center validation, enhancing explainable AI fused with medical knowledge, improving cross-device and cross-environment robustness through hardware-software co-design, refining human-AI collaborative diagnostic paradigms, and establishing comprehensive standardization and regulatory frameworks. Overcoming these hurdles through concerted interdisciplinary efforts is essential to realize a full-cycle intelligent health ecosystem, advancing precision medicine and equitable healthcare delivery.

Cite this article

Gao Yingming, Wu Yangqing, Yang Fei, Zhou Yingying, Li Ya, Wu Mengyue . Research Progress and Trends of Intelligent Speech in Pathological Healthcare[J]. Journal of Shanghai Jiaotong University(Science), 2026 , 31(3) : 671 -692 . DOI: 10.1007/s12204-026-2943-8

References

[1] Laennec R T H. De l’auscultation médiate, ou traité du diagnostic des maladies des poumons et du coeur [M]. Paris: J.-A. Brosson & J.-S. Chaudé, 1819.
[2] Fant G. Acoustic theory of speech production: With calculations based on X-ray studies of Russian articulation [M]. The Hague: Mouton, 1960.
[3] Wang Q Y, Fu Y, Shao B Y, et al. Early detection of Parkinson’s disease from multiple signal speech: Based on Mandarin language dataset [J]. Frontiers in Aging Neuroscience, 2022, 14: 1036588.
[4] Weizenbaum E L, Fulford D, Torous J, et al. Smartphone-based neuropsychological assessment in Parkinson’s disease: Feasibility, validity, and contextually driven variability in cognition [J]. Journal of the International Neuropsychological Society, 2022, 28(4): 401-413.
[5] Huang G, Li R J, Bai Q, et al. Multimodal learning of clinically accessible tests to aid diagnosis of neurodegenerative disorders: A scoping review [J]. Health Information Science and Systems, 2023, 11(1): 32.
[6] Bhattacharya D, Sharma N K, Dutta D, et al. Coswara: A respiratory sounds and symptoms dataset for remote screening of SARS-CoV-2 infection [J]. Scientific Data, 2023, 10(1): 397.
[7] Haider N S, Behera A K. Computerized lung sound based classification of asthma and chronic obstructive pulmonary disease (COPD) [J]. Biocybernetics and Biomedical Engineering, 2022, 42(1): 42-59.
[8] Rocha B M, Filos D, Mendes L, et al. Α respiratory sound database for the development of automated classification[M]//Precision medicine powered by pHealth and connected health. Singapore: Springer, 2018: 33-37.
[9] Rocha B M, Pessoa D, Marques A, et al. Automatic classification of adventitious respiratory sounds: A (un)solved problem? [J]. Sensors, 2021, 21(1): 57.
[10] Landry V, Christopoulos A, Guertin L, et al. Patterns of alaryngeal voice adoption and predictive factors of vocal rehabilitation failure following total laryngectomy [J]. Head & Neck, 2023, 45(10): 2657-2669.
[11] Zhang Y C, Zou X, Yang J S, et al. Multimodal laryngoscopic video analysis for assisted diagnosis of vocal fold paralysis [J]. Computer Speech & Language, 2026, 96: 101891.
[12] Shen Y, Yang H Y, Lin L. Automatic depression detection: An emotional audio-textual corpus and a GRU/BILSTM-based model [C]// 2022 IEEE International Conference on Acoustics, Speech and Signal Processing. Singapore: IEEE, 2022: 6247-6251.
[13] Zhao Q H, Geng S J, Wang B Y, et al. Deep learning in heart sound analysis: From techniques to clinical applications [J]. Health Data Science, 2024, 4: 182.
[14] Oliveira J, Renna F, Costa P D, et al. The CirCor DigiScope dataset: From murmur detection to murmur classification [J]. IEEE Journal of Biomedical and Health Informatics, 2022, 26(6): 2524-2535.
[15] Dong F Q, Qian K, Ren Z, et al. Machine listening for heart status monitoring: Introducing and benchmarking HSS: The heart sounds Shenzhen corpus [J]. IEEE Journal of Biomedical and Health Informatics, 2020, 24(7): 2082-2092.
[16] Alghifari M F, Gunawan T S, Kartiwi M. Development of sorrow analysis dataset for speech depression prediction [C]//2023 IEEE International Instrumentation and Measurement Technology Conference. Kuala Lumpur: IEEE, 2023: 1-6.
[17] Tasnim M, Ehghaghi M, Diep B, et al. DEPAC: A corpus for depression and anxiety detection from speech [C]// Eighth Workshop on Computational Linguistics and Clinical Psychology. Seattle: ACL, 2022: 1-16.
[18] Zou B C, Han J L, Wang Y X, et al. Semi-structural interview-based Chinese multimodal depression corpus towards automatic preliminary screening of depressive disorders [J]. IEEE Transactions on Affective Computing, 2023, 14(4): 2823-2838.
[19] Cai H S, Yuan Z Q, Gao Y W, et al. A multi-modal open dataset for mental-disorder analysis [J]. Scientific Data, 2022, 9: 178.
[20] Yoon J, Kang C, Kim S, et al. D-vlog: Multimodal vlog dataset for depression detection [C]//Proceedings of the Thirty-Sixth AAAI Conference on Artificial Intelligence. 2022, 36(11): 12226-12234.
[21] Tayebi Arasteh S, Weise T, Schuster M, et al. The effect of speech pathology on automatic speaker verification: A large-scale study [J]. Scientific Reports, 2023, 13: 20476.
[22] Tayebi Arasteh S, Arias-Vergara T, Pérez-Toro P A, et al. Addressing challenges in speaker anonymization to maintain utility while ensuring privacy of pathological speech [J]. Communications Medicine, 2024, 4: 182.
[23] Hsiao C H, Ruan S J, Chen C L, et al. A text-dependent end-to-end speech sound disorder detection and diagnosis in Mandarin-speaking children [J]. IEEE Transactions on Instrumentation and Measurement, 2024, 73: 3001911.
[24] Kojima T, Fujimura S, Hasebe K, et al. Objective assessment of pathological voice using artificial intelligence based on the GRBAS scale [J]. Journal of Voice, 2024, 38(3): 561-566.
[25] Liu Z Y, Li C C, Gao X, et al. Ensemble-based depression detection in speech [C]//2017 IEEE International Conference on Bioinformatics and Biomedicine. Kansas City: IEEE, 2017: 975-980.
[26] Scibelli F, Roffo G, Tayarani M, et al. Depression speaks: Automatic discrimination between depressed and non-depressed speakers based on nonverbal speech features [C]//2018 IEEE International Conference on Acoustics, Speech and Signal Processing. Calgary: IEEE, 2018: 6842-6846.
[27] Srimadhur N S, Lalitha S. An end-to-end model for detection and assessment of depression levels using speech [J]. Procedia Computer Science, 2020, 171: 12-21.
[28] Rusz J, Krack P, Tripoliti E. From prodromal stages to clinical trials: The promise of digital speech biomarkers in Parkinson’s disease [J]. Neuroscience & Biobehavioral Reviews, 2024, 167: 105922.
[29] Orozco-Arroyave J R, Arias-Londoño J D, Vargas-Bonilla J F, et al. New Spanish speech corpus database for the analysis of people suffering from Parkinson’s disease [C]// Ninth International Conference on Language Resources and Evaluation. Reykjavik: ACL, 2014: 342-347.
[30] Bowden M, Beswick E, Tam J, et al. A systematic review and narrative analysis of digital speech biomarkers in motor neuron disease [J]. npj Digital Medicine, 2023, 6: 228.
[31] Luz S, Haider F, de la Fuente S, et al. Alzheimer’s dementia recognition through spontaneous speech: The ADReSS challenge [C]//Interspeech 2020. Shanghai: ISCA, 2020: 2172-2176.
[32] Saeedi S, Hetjens S, Grimm M O W, et al. Acoustic speech analysis in Alzheimer’s disease: A systematic review and meta-analysis [J]. The Journal of Prevention of Alzheimer’s Disease, 2024, 11(6): 1789-1797.
[33] Sarkar M, Madabhavi I, Niranjan N, et al. Auscultation of the respiratory system [J]. Annals of Thoracic Medicine, 2015, 10(3): 158-168.
[34] Meslier N, Charbonneau G, Racineux J L. Wheezes [J]. The European Respiratory Journal, 1995, 8(11): 1942-1948.
[35] Vyshedskiy A, Alhashem R M, Paciej R, et al. Mechanism of inspiratory and expiratory crackles [J]. Chest, 2009, 135(1): 156-164.
[36] Kim Y, Hyon Y, Jung S S, et al. Respiratory sound classification for crackles, wheezes, and rhonchi in the clinical field using deep learning [J]. Scientific Reports, 2021, 11: 17186.
[37] Makimoto H, Shiraga T, Kohlmann B, et al. Efficient screening for severe aortic valve stenosis using understandable artificial intelligence: A prospective diagnostic accuracy study [J]. European Heart Journal - Digital Health, 2022, 3(2): 141-152.
[38] Chorba J S, Shapiro A M, Le L, et al. Deep learning algorithm for automated cardiac murmur detection via a digital stethoscope platform [J]. Journal of the American Heart Association, 2021, 10(9): e019905.
[39] Prince J, Maidens J, Kieu S, et al. Deep learning algorithms to detect murmurs associated with structural heart disease [J]. Journal of the American Heart Association, 2023, 12(20): e030377.
[40] Ainiwaer A, Hou W Q, Qi Q, et al. Deep learning of heart-sound signals for efficient prediction of obstructive coronary artery disease [J]. Heliyon, 2024, 10(1): e23354.
[41] Rabiner L R, Schafer R W. Digital processing of speech signals [M]. Englewood Cliffs: Prentice-Hall, 1978.
[42] Boersma P, Weenink D. Praat, a system for doing phonetics by computer [J]. Glot International, 2002, 5(9/10): 341-345.
[43] Shinde S G, Tambe A C, Vishwakarma A, et al. Automated depression detection using audio features [J]. International Research Journal of Engineering and Technology, 2020, 7(5): 1-5.
[44] Li Q F, Wang D, Ren Y M, et al. FTA-net: A frequency and time attention network for speech depression detection [C]// Interspeech 2023. Dublin: ISCA, 2023: 1723-1727.
[45] Verma A, Jain P, Kumar T. An effective depression diagnostic system using speech signal analysis through deep learning methods [J]. International Journal on Artificial Intelligence Tools, 2023, 32(2): 2340004.
[46] Sun C J, Jiang M, Gao L L, et al. A novel study for depression detecting using audio signals based on graph neural network [J]. Biomedical Signal Processing and Control, 2024, 88: 105675.
[47] Yue X P, Zhang C N, Wang Z J, et al. Hierarchical transformer speech depression detection model research based on dynamic window and attention merge [J]. PeerJ Computer Science, 2024, 10: e2348.
[48] Gupta S, Agarwal G, Agarwal S, et al. Depression detection using cascaded attention based deep learning framework using speech data [J]. Multimedia Tools and Applications, 2024, 83(25): 66135-66173.
[49] Wu W, Wu M Y, Yu K. Climate and weather: Inspecting depression detection via emotion recognition [C]// 2022 IEEE International Conference on Acoustics, Speech and Signal Processing. Singapore: IEEE, 2022: 6262-6266.
[50] Zhang P Y, Wu M Y, Dinkel H, et al. DEPA: Self-supervised audio embedding for depression detection [C]// 29th ACM International Conference on Multimedia. Online: ACM, 2021: 135-143.
[51] Hsu W N, Bolte B, Tsai Y H, et al. HuBERT: Self-supervised speech representation learning by masked prediction of hidden units [J]. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2021, 29: 3451-3460.
[52] Radford A, Kim J W, Xu T, et al. Robust speech recognition via large-scale weak supervision [C]//Proceedings of the 40th International Conference on Machine Learning. Honolulu: PMLR, 2023: 28492-28518.
[53] Li S T, Chu X, Zhang Z Y, et al. Study on depression detection model with self-supervised speech representation: The architecture of wav2vec2-LSTM-CNN [C]//2025 IEEE 7th International Conference on Communications, Information System and Computer Engineering. Guangzhou: IEEE, 2025: 1195-1201.
[54] Zhang P Y, Wu M Y, Yu K. ReCLR: Reference-enhanced contrastive learning of audio representation for depression detection [C]//Interspeech 2023. Dublin: ISCA, 2023: 2998-3002.
[55] Zhang X Y, Liu H X, Xu K S, et al. When LLMs meets acoustic landmarks: An efficient approach to integrate speech into large language models for depression detection [C]//Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Miami: ACL, 2024: 146-158.
[56] Teng S Y, Liu J Q, Sun H, et al. Enhanced multimodal depression detection with emotion prompts [C]// 2025 IEEE International Conference on Acoustics, Speech and Signal Processing. Hyderabad: IEEE, 2025: 1-5.
[57] Lin Z L, Wang Y W, Zhou Y J, et al. MLM-EOE: Automatic depression detection via sentimental annotation and multi-expert ensemble [J]. IEEE Transactions on Affective Computing, 2025, 16(4): 2842-2858.
[58] Verma S R, Shevatkar S V, Bhange S S, et al. Mental health support chatbot with AI counselling [J]. International Journal for Multidisciplinary Research, 2025, 7: 37678.
[59] Saeed M A, Komashinsky V, Mohammed S A, et al. Speech signal analysis to predict depression [J]. International Journal of Advanced Networking and Applications, 2025, 16(4): 6460-6465.
[60] Abubakar A M, Gupta D, Parida S. A reinforcement learning approach for intelligent conversational chatbot for enhancing mental health therapy [J]. Procedia Computer Science, 2024, 235: 916-925.
[61] Gao Y. The impact and application of artificial intelligence technology on mental health counseling services for college students [J]. Journal of Computational Methods in Sciences and Engineering, 2024: 1-18.
[62] Shankar Ganesh M, Venkateswaramurthy N. Artificial intelligence (AI) generated health counseling for mental illness patients [J]. Current Psychiatry Research and Reviews, 2025, 21(3): 269-283.
[63] Nurfadhilah E, Yoganingrum A, Latief A D, et al. Artificial intelligence technologies in mental health: Transforming depression care through innovation [M]//Humanizing technology with emotional intelligence. Hershey: IGI Global, 2024: 219-262.
[64] Parimala M. Soul support: AI driven emotional assistance ChatBot [J]. International Journal of Scientific Research in Engineering and Management, 2025, 9(6): 1-9.
[65] Sharma A, Saxena A, Kumar A, et al. Depression detection using multimodal analysis with chatbot support [C]//2024 2nd International Conference on Disruptive Technologies. Greater Noida: IEEE, 2024: 328-334.
[66] Naidu S M M, Chavan Y A, Pillai M R, et al. Detection of depression and anxiety through speech, voice, and sentiment analysis [J]. The Ciência & Engenharia - Science & Engineering Journal, 2023, 11(1): 703-711.
[67] Zhang Y, Huang B. Design and implementation of speech emotion interaction system for patients with mental illness [J]. Schizophrenia Bulletin, 2025, 51(Supplement_1): S6-S7.
[68] Moghe B, Kachhara M, M K. Multimodal emotion analysis for depression detection- integrating facial expression and speech recognition [C]//2024 Second International Conference on Inventive Computing and Informatics. Bangalore: IEEE, 2024: 37-42.
[69] Jadhav M S, Kushwaha E, Tripathy A, et al. Emotion aware AI for mental health monitoring [J]. International Journal of Advanced Research in Science, Communication and Technology, 2024: 63-69.
[70] Siddals S, Torous J, Coxon A. “It happened to be the perfect thing”: Experiences of generative AI chatbots for mental health [J]. npj Mental Health Research, 2024, 3: 48.
[71] Thieme A, Hanratty M, Lyons M, et al. Designing human-centered AI for mental health: Developing clinically relevant applications for online CBT treatment [J]. ACM Transactions on Computer-Human Interaction, 2023, 30(2): 1-50.
[72] Zhong Z, Wang Z. Intelligent depression prevention via LLM-based dialogue analysis: Overcoming the limitations of scale-dependent diagnosis through precise emotional pattern recognition [PP/OL]. arXiv (2025-04-23) [2026-01-18]. https://doi.org/10.48550/arXiv.2504.16504.

Outlines

/