Celebrating 30 Years

Hybrid Supervised Fine-Tuning Method for Medical Language Models via Explicit Reasoning Modeling

Expand
  • 1. School of Automation and Intelligent Sensing, Shanghai Jiao Tong University, Shanghai 200240, China; 2. Shanghai West Hongqiao Navigation Technology Co., Ltd., Shanghai 201702, China

Received date: 2025-12-09

  Revised date: 2026-01-20

  Accepted date: 2026-02-25

  Online published: 2026-05-28

Abstract

To enhance the reasoning stability of small-parameter medical large language models in internal medicine question-answering tasks, this paper proposes a training methodology based on explicit chain-of-thought (CoT) modeling and hybrid supervised fine-tuning (SFT). First, a hierarchical dataset comprising general internal medicine instructions and explicit CoT data was constructed. On this basis, a two-stage hybrid SFT process was implemented, incorporating direct preference optimization to align the model with clinical preferences. Experimental results demonstrate that the proposed method improves the accuracy on Chinese medical benchmarks while reducing the proportion of redundant reasoning, effectively enhancing the logical rigor of complex clinical inquiries. Furthermore, these findings validate the potential of this approach for deploying low-cost, highly reliable, and localized auxiliary diagnostic systems in privacy-sensitive and compute-constrained clinical scenarios.

Cite this article

Wang Xu, Tao Wei, Nan Zhuojiang, Wan Song . Hybrid Supervised Fine-Tuning Method for Medical Language Models via Explicit Reasoning Modeling[J]. Journal of Shanghai Jiaotong University(Science), 2026 , 31(3) : 660 -670 . DOI: 10.1007/s12204-026-2929-6

References

[1] Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need [C]// 31st Conference on Neural Information Processing Systems. Long Beach: NIPS, 2017: 5998-6008.
[2] Touvron H, Martin L, Stone K, et al. Llama 2: Open foundation and fine-tuned chat models [PP/OL]. V2. arXiv (2023-07-19) [2025-12-06]. https://doi.org/10.48550/arXiv.2307.09288.
[3] Zhao L L, Zeng W H, Shi X F, et al. CareBot: A pioneering full-process open-source medical language model [J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2025, 39(24): 26039-26047.
[4] Garg M, Raza S, Rayana S, et al. The rise of small language models in healthcare: A comprehensive survey [PP/OL]. V2. arXiv (2025-04-25) [2025-12-06]. https://doi.org/10.48550/arXiv.2504.17119.
[5] Lin T W, Zhang W Q, Li S J, et al. HealthGPT: A medical large vision-language model for unifying comprehension and generation via heterogeneous knowledge adaptation [PP/OL]. V3. arXiv (2025-02-21) [2025-12-06]. https://doi.org/10.48550/arXiv.2502.09838.
[6] Wu C Y, Lin W X, Zhang X M, et al. PMC-LLaMA: Toward building open-source language models for medicine [J]. Journal of the American Medical Informatics Association, 2024, 31(9): 1833-1843.
[7] Wang S S, Hu M Z, Li Q, et al. Capabilities of GPT-5 on multimodal medical reasoning [PP/OL]. V2. arXiv (2025-08-13) [2025-12-06]. https://doi.org/10.48550/arXiv.2508.08224.
[8] Comanici G, Bieber E, Schaekermann M, et al. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities [PP/OL]. arXiv (2025-07-07) [2025-12-06]. https://doi.org/10.48550/arXiv.2507.06261.
[9] Huang L, Yu W J, Ma W T, et al. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions [J]. ACM Transactions on Information Systems, 2025, 43(2): 1-55.
[10] Nye M, Andreassen A J, Gur-Ari G, et al. Show your work: Scratchpads for intermediate computation with language models [PP/OL]. arXiv (2021-11-30) [2025-12-05]. https://doi.org/10.48550/arXiv.2112.00114.
[11] Wei J, Wang X Z, Schuurmans D, et al. Chain-of-thought prompting elicits reasoning in large language models [PP/OL]. V6. arXiv (2023-01-10) [2025-12-05]. https://doi.org/10.48550/arXiv.2201.11903.
[12] Wu Z L, Hasan A, Wu J G, et al. Chain-of-Though (CoT) prompting strategies for medical error detection and correction [PP/OL]. arXiv (2024-06-13) [2025-12-05]. https://doi.org/10.48550/arXiv.2406.09103.
[13] Wang X Z, Wei J, Schuurmans D, et al. Self-consistency improves chain of thought reasoning in language models [PP/OL]. V4. arXiv (2023-03-07) [2025-12-05]. https://doi.org/10.48550/arXiv.2203.11171.
[14] Ton J F, Taufiq M F, Liu Y. Understanding chain-of-thought in LLMs through information theory [PP/OL]. V2. arXiv (2025-07-10) [2025-12-05]. https://doi.org/10.48550/arXiv.2411.11984.
[15] Merrill W, Sabharwal A. The expressive power of transformers with chain of thought [PP/OL]. V5. arXiv (2024-04-11) [2025-12-05]. https://doi.org/10.48550/arXiv.2310.07923.
[16] Ye Y X, Huang Z, Xiao Y, et al. LIMO: Less is more for reasoning [PP/OL]. V3. arXiv (2025-07-29) [2025-12-06]. https://doi.org/10.48550/arXiv.2502.03387.
[17] Wang X, Chen G, Dingjie S, et al. CMB: A comprehensive medical benchmark in Chinese [C]// 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Mexico City: ACL, 2024: 6184-6205. 
[18] Liu J L, Zhou P L, Hua Y N, et al. Benchmarking large language models on CMExam: A comprehensive Chinese medical exam dataset [PP/OL]. V3. arXiv (2023-10-23) [2025-12-05]. https://doi.org/10.48550/arXiv.2306.03030.
[19] Li H, Zhang Y, Koto F, et al. CMMLU: Measuring massive multitask language understanding in Chinese [C]//Findings of the Association for Computational Linguistics. Bangkok: ACL, 2024: 11260-11285. 
[20] Rafailov R, Sharma A, Mitchell E, et al. Direct preference optimization: Your language model is secretly a reward model [PP/OL]. V3. arXiv (2024-07-29) [2025-12-05]. https://doi.org/10.48550/arXiv.2305.18290.
[21] Huang Y, Bai Y, Zhu Z, et al. C-Eval: A multi-level multi-discipline Chinese evaluation suite for foundation models [C]// 37th Conference on Neural Information Processing Systems. New Orleans: NIPS, 2023: 62991-63010. 
[22] Kwon W, Li Z, Zhuang S, et al. Efficient memory management for large language model serving with PagedAttention [C]// 29th Symposium on Operating Systems Principles. Koblenz: ACM, 2023: 611-626.
[23] Lewis P, Perez E, Piktus A, et al. Retrieval-augmented generation for knowledge-intensive NLP tasks [C]// 34th Conference on Neural Information Processing Systems. Vancouver: NIPS, 2020: 9459-9474.

Outlines

/