This paper presents a costmap A* guided soft actor-critic (CMA-SAC) path planning method to optimize the navigation performance of robots in long-distance and complex environments. Initially, a costmap is constructed to calculate the cost for approaching obstacles. With the costmap, the improved A* algorithm effectively avoids the paths being too close to obstacles. Subsequently, a local path planner based on deep reinforcement learning is constructed to directly generate control commands for the robot. Lastly, a tightly coupled strategy of global and local path planning is employed, where the results of global path planning are incorporated as part of the input to the deep neural network and integrated into the reward function of reinforcement learning (RL). Simulation experiments indicate that the CMA-SAC method outperforms deep deterministic policy gradient and SAC algorithms in terms of learning speed and stability during training. And in the test tasks, the CMA-SAC method performs better than other RL-based methods in navigation efficiency and has better dynamic obstacle avoidance performance than dynamic window approach. The proposed method has a success rate of 95.8% in the maze environment and the highest success rate in the long-distance and dynamic environment, demonstrating the method’s ability in complex navigation tasks.
Wang Yixuan, Shen Bin, Xiong Leilei, Nan Zhuojiang, Tao Wei
. Costmap A∗ Guided Reinforcement Learning Path Planning Method for Complex Environments Navigation[J]. Journal of Shanghai Jiaotong University(Science), 2026
, 31(4)
: 819
-827
.
DOI: 10.1007/s12204-024-2755-7
[1] SAHOO S K, CHOUDHURY B B. A review of methodologies for path planning and optimization of mobile robots [J]. Journal of Process Management and New Technologies, 2023, 11(1/2): 122-140.
[2] RAFAI A N A, ADZHAR N, JAINI N I. A review on path planning and obstacle avoidance algorithms for autonomous mobile robots [J]. Journal of Robotics, 2022, 2022: 2538220.
[3] QIN H W, SHAO S L, WANG T, et al. Review of autonomous path planning algorithms for mobile robots [J]. Drones, 2023, 7(3): 211.
[4] LIU Yahui, SHEN Xingwang, GU Xinghai, et al. A dual-system reinforcement learning method for flexible job shop dynamic scheduling [J]. Journal of Shanghai Jiao Tong University, 2022, 56(9): 1262-1275 (in Chinese).
[5] ALMAZROUEI K, KAMEL I, RABIE T. Dynamic obstacle avoidance and path planning through reinforcement learning [J]. Applied Sciences, 2023, 13(14): 8174.
[6] YIN Y, CHEN Z Y, LIU G, et al. A mapless local path planning approach using deep reinforcement learning framework [J]. Sensors, 2023, 23(4): 2036.
[7] FU G, GAO Y, LIU L W, et al. UAV mission path planning based on reinforcement learning in dynamic environment [J]. Journal of Function Spaces, 2023, 2023: 9708143.
[8] KHLIF N, NAHLA K, SAFYA B. Reinforcement learning with modified exploration strategy for mobile robot path planning [J]. Robotica, 2023, 41(9): 2688-2702.
[9] ZHANG K, HU Y J, HUANG D Q, et al. Target tracking and path planning of mobile sensor based on deep reinforcement learning [C]//2023 IEEE 12th Data Driven Control and Learning Systems Conference. Xiangtan: IEEE, 2023: 190-195.
[10] SHI Z, WANG K Y, ZHANG J H. Improved reinforcement learning path planning algorithm integrating prior knowledge [J]. PLoS One, 2023, 18(5): e0284942.
[11] TU G T, JUANG J G. UAV path planning and obstacle avoidance based on reinforcement learning in 3D environments [J]. Actuators, 2023, 12(2): 57.
[12] YIN Z Q, CAO W, SONG T, et al. Reinforcement learning path planning based on step batch Q-learning algorithm [C]//2022 IEEE International Conference on Artificial Intelligence and Computer Applications. Dalian: IEEE, 2022: 630-633.
[13] YU X B, LUO W G. Reinforcement learning-based multi-strategy cuckoo search algorithm for 3D UAV path planning [J]. Expert Systems with Applications, 2023, 223: 119910.
[14] YANG J C, NI J F, XI M, et al. Intelligent path planning of underwater robot based on reinforcement learning [J]. IEEE Transactions on Automation Science and Engineering, 2023, 20(3): 1983-1996.
[15] CHEN L, WANG Y N, MIAO Z Q, et al. Transformer-based imitative reinforcement learning for multirobot path planning [J]. IEEE Transactions on Industrial Informatics, 2023, 19(10): 10233-10243.
[16] HUANG W H, ZHOU Y X, HE X K, et al. Goal-guided transformer-enabled reinforcement learning for efficient autonomous navigation [J]. IEEE Transactions on Intelligent Transportation Systems, 2024, 25(2): 1832-1845.
[17] LEE M H, MOON J. Deep reinforcement learning-based model-free path planning and collision avoidance for UAVs: A soft actor–critic with hindsight experience replay approach [J]. ICT Express, 2023, 9(3): 403-408.
[18] JI X K, HAI J T, LUO W G, et al. Obstacle avoidance in multi-agent formation process based on deep reinforcement learning [J]. Journal of Shanghai Jiao Tong University (Science), 2021, 26(5): 680-685.
[19] JI M Y, LI J X, LI S Q, et al. Research on path planning of mobile robot based on reinforcement learning [C]//2022 China Automation Congress. Xiamen: IEEE, 2022: 748-751.
[20] ZHU K, ZHANG T. Deep reinforcement learning based mobile robot navigation: A review [J]. Tsinghua Science and Technology, 2021, 26(5): 674-691.
[21] VIDYASAGAR M. A tutorial introduction to reinforcement learning [J]. SICE Journal of Control, Measurement, and System Integration, 2023, 16(1): 172-191.
[22] HART P E, NILSSON N J, RAPHAEL B. A formal basis for the heuristic determination of minimum cost paths [J]. IEEE Transactions on Systems Science and Cybernetics, 1968, 4(2): 100-107.