J Shanghai Jiaotong Univ Sci ›› 2026, Vol. 31 ›› Issue (1): 154-166.doi: 10.1007/s12204-025-2851-3
收稿日期:2024-12-16
修回日期:2025-02-25
接受日期:2025-03-21
出版日期:2026-02-28
发布日期:2025-10-14
夏洁1,吴晓东1,许敏2
Received:2024-12-16
Revised:2025-02-25
Accepted:2025-03-21
Online:2026-02-28
Published:2025-10-14
摘要: 端到端自动驾驶技术打破了传统自动驾驶技术模块化管道形式的桎梏,将感知、预测、规划集成在一个框架下,实现了全局优化。目前较为典型的端到端(E2E)框架都是基于深度学习规划的,需要大量真实世界的离线数据对网络进行训练,而数据的获取与管理是一件费时又费力的事情。基于深度强化学习(DRL)算法进行规划也是当今流行的一种自动驾驶技术,深度强化学习算法能够促使智能体在环境突变时通过奖励函数的引导实现自适应,但这类学习框架与感知模块之间没有强关联性,即无法实现反向传播。上述的两类学习框架各有优缺点,本文选择将两个框架融合在一起,并搭建了鸟瞰图(BEV)特征提取网络从相机拍摄的图像中提取关键交通流信息,最终构建出基于BEV特征的端到端深度强化学习规划框架,该框架使得端到端自动驾驶技术由数据驱动转化为行为驱动。为了提高网络训练速度与质量,本文还提出了先进的模仿学习算法。所提出的算法最后在CARLA仿真器中进行仿真验证,实验结果证明该算法优于其他框架下的算法,能够进一步提高智能体的安全性、高效性等。
中图分类号:
. 融合鸟瞰图特征的模仿与强化学习自动驾驶规划方法[J]. J Shanghai Jiaotong Univ Sci, 2026, 31(1): 154-166.
Xia Jie, Wu Xiaodong, Xu Min. BEV-Fused Imitation and Reinforcement Learning for Autonomous Driving Planning[J]. J Shanghai Jiaotong Univ Sci, 2026, 31(1): 154-166.
| [1] CHEN L, WU P H, CHITTA K, et al. End-to-end autonomous driving: Challenges and frontiers [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46(12): 10164-10183. [2] HU S C, CHEN L, WU P H, et al. ST-P3: End-to-end vision-based autonomous driving viaSpatial-temporal feature learning [M]//Computer Vision – ECCV 2022. Cham: Springer, 2022: 533-549. [3] HU Y H, YANG J Z, CHEN L, et al. Planning-oriented autonomous driving [C]//2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023: 17853-17862. [4] YE T, JING W, HU C, et al. FusionAD: Multi-modality fusion for prediction and planning tasks of autonomous driving [DB/OL]. (2023-08-02). https://arxiv.org/abs/2308.01006 [5] MUHAMMAD K, ULLAH A, LLORET J, et al. Deep learning for safe autonomous driving: Current challenges and future directions [J]. IEEE Transactions on Intelligent Transportation Systems, 2021, 22(7): 4316-4336. [6] BASILE G, LECCESE S, PETRILLO A, et al. Sustainable DDPG-based path tracking for connected autonomous electric vehicles in extra-urban scenarios [J]. IEEE Transactions on Industry Applications, 2024, 60(6): 9237-9250. [7] REN Y G, DUAN J L, LI S E, et al. Improving generalization of reinforcement learning with minimax distributional soft actor-critic [C]//2020 IEEE 23rd International Conference on Intelligent Transportation Systems. Rhodes: IEEE, 2020: 1-6. [8] LI S Y, LI M Z, JING Z L. Multi-agent path planning method based on improved deep Q-network in dynamic environments [J]. Journal of Shanghai Jiao Tong University (Science), 2024, 29(4): 601-612. [9] WU J D, HUANG Z Y, HANG P, et al. Digital twin-enabled reinforcement learning for end-to-end autonomous driving [C]//2021 IEEE 1st International Conference on Digital Twins and Parallel Intelligence. Beijing: IEEE, 2021: 62-65. [10] HUANG Z Q, ZHANG J, TIAN R, et al. End-to-end autonomous driving decision based on deep reinforcement learning [C]//2019 5th International Conference on Control, Automation and Robotics. Beijing: IEEE, 2019: 658-662. [11] LI H Y, SIMA C, DAI J F, et al. Delving into the Devils of bird’s-eye-view perception: A review, evaluation and recipe [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46(4): 2151-2170. [12] MA Y X, WANG T, BAI X Y, et al. Vision-centric BEV perception: A survey [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024, 46(12): 10978-10997. [13] LI Z Q, WANG W H, LI H Y, et al. BEVFormer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers [M]//Computer Vision – ECCV 2022. Cham: Springer, 2022: 1-18. [14] LIU Z J, TANG H T, AMINI A, et al. BEVFusion: Multi-task multi-sensor fusion with unified bird’s-eye view representation [C]//2023 IEEE International Conference on Robotics and Automation. London: IEEE, 2023: 2774-2781. [15] RAMRAKHYA R, BATRA D, WIJMANS E, et al. PIRLNav: Pretraining with imitation and RL finetuning for OBJECTNAV [C]//2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition. Vancouver: IEEE, 2023: 17896-17906. [16] Y G D, NAIR N G, SATPATHY P, et al. Covariate shift: A review and analysis on classifiers [C]//2019 Global Conference for Advancement in Technology. Bangalore: IEEE, 2019: 1-6. [17] KURNIAWATI H. Partially observable Markov decision processes (POMDPs) and robotics [DB/OL]. (2021-07-15). https://arxiv.org/abs/2107.07599 [18] HUBMANN C, SCHULZ J, BECKER M, et al. Automated driving in uncertain environments: Planning with interaction and uncertain maneuver prediction [J]. IEEE Transactions on Intelligent Vehicles, 2018, 3(1): 5-17. [19] DOSOVITSKIY A, ROS G, CODEVILLA F, et al. CARLA: An open urban driving simulator [C]// 1st Annual Conference on Robot Learning. Mountain View: PMLR, 2017: 1-16. [20] LI Y. Deep reinforcement learning: An overview [DB/OL]. (2017-01-25). https://arxiv.org/abs/1701.07274 [21] PHILION J, FIDLER S. Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3D [M]//Computer Vision – ECCV 2020. Cham: Springer, 2020: 194-210. [22] HUANG J, HUANG G, ZHU Z, et al. BEVDet: High-performance multi-camera 3D object detection in bird-eye-view [DB/OL]. (2021-12-22). https://arxiv.org/abs/2112.11790 [23] CHEN X Z, MA H M, WAN J, et al. Multi-view 3D object detection network for autonomous driving [C]//2017 IEEE Conference on Computer Vision and Pattern Recognition. Honolulu: IEEE, 2017: 6526-6534. [24] O'SHEA K, NASH R. An introduction to convolutional neural networks [DB/OL]. (2015-11-26). https://arxiv.org/abs/1511.08458 [25] SCHULMAN J, WOLSKI F, DHARIWAL P, et al. Proximal policy optimization algorithms [DB/OL]. (2017-07-20). https://arxiv.org/abs/1707.06347 [26] SCHULMAN J, MORITZ P, LEVINE S, et al. High-dimensional continuous control using generalized advantage estimation [DB/OL]. (2015-06-08). https://arxiv.org/abs/1506.02438 [27] ZARE M, KEBRIA P M, KHOSRAVI A, et al. A survey of imitation learning: Algorithms, recent developments, and challenges [J]. IEEE Transactions on Cybernetics, 2024, 54(12): 7173-7186. [28] ZHANG Z J, LINIGER A, DAI D X, et al. End-to-end urban driving by imitating a reinforcement learning coach [C]//2021 IEEE/CVF International Conference on Computer Vision. Montreal: IEEE, 2021: 15202-15212. |
| [1] | 袁景美, 赵亮, 孙卓然, 徐志朝, 牛亚雷. 基于深度强化学习的导航信号自适应干扰决策方法[J]. 空天防御, 2026, 9(2): 41-52. |
| [2] | . 触觉辅助导航车辆:增强盲区和透明物体场景中的障碍物检测[J]. J Shanghai Jiaotong Univ Sci, 2026, 31(1): 167-175. |
| [3] | 陈实, 杨林森, 刘艺洪, 罗欢, 臧天磊, 周步祥. 小样本数据驱动模式下的新建微电网优化调度策略[J]. 上海交通大学学报, 2025, 59(6): 732-745. |
| [4] | 王志博, 呼卫军, 马先龙, 全家乐, 周皓宇. 感知驱动控制的无人机拦截碰撞技术[J]. 空天防御, 2025, 8(4): 78-84. |
| [5] | 周文杰, 付昱龙, 郭相科, 戚玉涛, 张海宾. 基于博弈树与数字平行战场的空战决策方法[J]. 空天防御, 2025, 8(3): 50-58. |
| [6] | 刘雁行, 乔如妤, 梁楠, 陈宇, 于凯, 吴汉霄. 基于负荷准线和深度强化学习的含电动汽车集群系统新能源消纳策略[J]. 上海交通大学学报, 2025, 59(10): 1464-1475. |
| [7] | 杨映荷, 魏汉迪, 范迪夏, 李昂. 基于高斯过程回归和深度强化学习的水下扑翼推进性能寻优方法[J]. 上海交通大学学报, 2025, 59(1): 70-78. |
| [8] | 周毅, 周良才, 史迪, 赵小英, 闪鑫. 基于安全深度强化学习的电网有功频率协同优化控制[J]. 上海交通大学学报, 2024, 58(5): 682-692. |
| [9] | 董玉博1, 崔涛1, 周禹帆1, 宋勋2, 祝月2, 董鹏1. 基于长周期极坐标系追击问题的多智能体强化学习奖赏函数设计方法[J]. J Shanghai Jiaotong Univ Sci, 2024, 29(4): 646-655. |
| [10] | 李舒逸, 李旻哲, 敬忠良. 动态环境下基于改进DQN的多智能体路径规划方法[J]. J Shanghai Jiaotong Univ Sci, 2024, 29(4): 601-612. |
| [11] | 苗镇华1, 黄文焘2, 张依恋3, 范勤勤1. 基于深度强化学习的多模态多目标多机器人任务分配算法[J]. J Shanghai Jiaotong Univ Sci, 2024, 29(3): 377-387. |
| [12] | 全家乐, 马先龙, 沈昱恒. 基于近端策略动态优化的多智能体编队方法[J]. 空天防御, 2024, 7(2): 52-62. |
| [13] | 张威振, 何真, 汤张帆. 风扰下无人机栖落机动的强化学习控制设计[J]. 上海交通大学学报, 2024, 58(11): 1753-1761. |
| [14] | 马驰, 张国群, 孙俊格, 吕广喆, 张涛. 基于深度强化学习的综合电子系统重构方法[J]. 空天防御, 2024, 7(1): 63-70. |
| [15] | 李鹏, 阮晓钢, 朱晓庆, 柴洁, 任顶奇, 刘鹏飞. 基于深度强化学习的区域化视觉导航方法[J]. 上海交通大学学报, 2021, 55(5): 575-585. |
| 阅读次数 | ||||||
|
全文 |
|
|||||
|
摘要 |
|
|||||