融合MPC参数化控制与改进SAC的全向移动机器人能耗优化与路径规划
DOI:
CSTR:
作者:
作者单位:

1.长春理工大学电子信息工程学院长春130022; 2.长春理工大学吉林省智能机器人高校协同 创新中心长春130022; 3.长春汽车职业技术大学电气工程学院长春130012

作者简介:

通讯作者:

中图分类号:

TH134TP242

基金项目:

吉林省科技发展计划项目(YDZJ202503CGZH002)资助


Energy consumption optimization and path planning of omnidirectional mobile robots integrating MPC parameterized control and improved SAC
Author:
Affiliation:

1.School of Electronic and Information Engineering, Changchun University of Science and Technology, Changchun 130022, China; 2.Jilin Provincial Collaborative Innovation Center for Intelligent Robots, Changchun University of Science and Technology, Changchun 130022, China; 3.College of Electrical Engineering, Changchun Vocational University of Automobile Technology, Changchun 130012, China

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    针对全向移动机器人在多障碍物不确定环境下的节能路径规划问题,提出一种考虑能耗的融合模型预测控制参数化与改进的软演员-评论家(SAC)框架的路径规划(E-QPSAC)算法。首先,将局部路径规划构造为二次规划问题,并采用改进原始-对偶混合梯度算法高效求解;其次,构建了SAC的12维状态空间(包括机器人位姿、目标信息及障碍物信息)、级联式网络架构及动态权重融合机制,有效平衡避障与目标趋向;同时考虑全向轮移动机器人因环境不确定性、系统非线性及运动特性导致的高能耗的特性,建立五维能耗模型,并依据能耗模型设计节能奖励函数。最后,在仿真平台中,与SAC、深度确定性策略梯度(DDPG)及基于优先经验回放与长短期记忆网络的双延迟深度确定性策略梯度(PL-TD3)算法进行对比实验。实验结果表明:所提算法任务完成率达98.7%(SAC为85.0%、DDPG为98.5%、PL-TD3为98.2%),总能耗低至18.274 kJ(SAC为25.765 kJ、DDPG为22.418 kJ、PL-TD3为21.191 kJ),平均路径长度6.45 m(SAC为6.56 m、DDPG为6.97 m、PL-TD3为7.13 m),所提算法在任务完成率、能耗、路径长度及鲁棒性上均更优。最终通过实测实验验证了所提方法的有效性。

    Abstract:

    To address the energy-efficient path planning problem of omnidirectional mobile robots in uncertain environments with multiple obstacles, this paper proposes an energy-aware path planning algorithm named energy-aware quadratic programming parameterization soft actor-critic(E-QPSAC), which integrates model predictive control parameterization with an improved soft actor-critic (SAC) framework. First, local path planning is formulated as a Quadratic Programming problem and efficiently solved by an improved primal-dual hybrid gradient algorithm. Second, a 12-dimensional state space (including robot pose, target information, and obstacle information) for SAC is constructed, along with a cascaded network architecture and a dynamic weight fusion mechanism to effectively balance obstacle avoidance and goal orientation. Meanwhile, considering the high energy consumption characteristics of omnidirectional mobile robots caused by environmental uncertainty, system nonlinearity, and motion characteristics, a five-dimensional energy consumption model is established, and an energy-saving reward function is designed based on this model. Finally, comparative experiments are conducted on a simulation platform with SAC, deep deterministic policy gradient (DDPG), and the prioritized replay and LSTM-based twin delayed deep deterministic policy gradient (PL-TD3) algorithm. The experimental results show that the proposed algorithm achieves a task completion rate of 98.7% (85.0% for SAC, 98.5% for DDPG, and 98.2% for PL-TD3), a total energy consumption as low as 18.274 kJ (25.765 kJ for SAC, 22.418 kJ for DDPG, and 21.191 kJ for PL-TD3), and an average path length of 6.45 m (6.56 m for SAC, 6.97 m for DDPG, and 7.13 m for PL-TD3). The proposed algorithm outperforms the others in task completion rate, energy consumption, path length, and robustness. The effectiveness of the proposed method is finally verified through practical experiments.

    参考文献
    相似文献
    引证文献
引用本文

单泽彪,何立本,刘小松,郑海平,梁法辉.融合MPC参数化控制与改进SAC的全向移动机器人能耗优化与路径规划[J].仪器仪表学报,2026,47(5):329-338

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:
  • 最后修改日期:
  • 录用日期:
  • 在线发布日期: 2026-07-24
  • 出版日期:
文章二维码