基于深度强化学习的仿蟹机器人运动控制研究
DOI:
CSTR:
作者:
作者单位:

东南大学仪器科学与工程学院南京210096

作者简介:

通讯作者:

中图分类号:

TH113TP242 TP181

基金项目:

国家自然科学基金项目(62173090)资助


Research on the deep reinforcement learning based locomotion control of a crab-like robot
Author:
Affiliation:

School of Instrument Science and Engineering, Southeast University, Nanjing 210096, China

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    当前双足和四足机器人的强化学习运动控制研究已取得显著进展,但八足机器人由于自由度冗余高、足端易干涉和动力学建模复杂等问题,其深度强化学习控制仍处于探索阶段。针对这一问题,提出了一种基于深度强化学习的仿蟹八足机器人运动控制框架。首先,通过分析螃蟹的身体结构和运动步态,设计了仿蟹机器人的机构模型和参考步态,为强化学习提供物理先验知识。随后,设计了模仿螃蟹步态参考动作的端到端深度强化学习训练框架,并引入行为克隆初始化策略以加速策略收敛,同时采用分阶段强化学习方法逐步增加地形难度,保证机器人在平地步态基础上稳定迁移至复杂地形。进一步对比分析了不同学习率调整策略、近端策略优化算法(provimal policy optimization,PPO)与优势演员-评论家(A2C)、软演员-评论家(SAC)、深度确定性策略梯度(DDPG)算法的适配性,并通过域随机化训练及行为克隆+分阶段强化学习方法研究了模型在动力学扰动和多样地形下的泛化能力。仿真实验结果显示,所提出框架在训练效率上优于其他算法,步态稳定性指标(即五步质心高度标准差)约为0.004 m;机器人在斜坡、台阶等复杂地形中均能实现稳定行进,运动性能优于训练的模型。最后,进行了初步样机实验,验证了所提方法在真实物理环境中的可行性和稳定性。所提出的深度强化学习框架结合了生物步态模仿、行为克隆初始化及分阶段强化学习,对八足机器人的运动控制提供了新的技术思路。

    Abstract:

    The research of reinforcement learning-based motion control has progressed significantly for the bipedal and quadrupedal robots. However, the deep reinforcement learning control of octopod robots is still in the exploratory stage due to the issues such as high degrees of freedom redundancy, easy interference at the foot end, and complex dynamic modeling. To address this problem, this paper proposes a motion control framework of crab-like octopod robots based on the deep reinforcement learning. Firstly,, the mechanism model and reference gait of the crab-like robot are designed by analyzing the body structure and gait of crabs, which provides the physical prior knowledge for reinforcement learning. Subsequently, an end-to-end deep reinforcement learning training framework is designed to imitate the reference actions of crab gaits, where a behavior cloning initialization strategy is introduced to accelerate the convergence of the policy. At the same time, a phased reinforcement learning method is adopted to gradually increase the terrain difficulty, ensuring that the robot can stably transfer from the flat ground gait to complex terrains. This paper further compares and analyzes different learning rate adjustment strategies as well as the adaptability of the proximal policy optimization (PPO) algorithm with actor-critic (A2C), soft actor critic (SAC), and deep deterministic policy gradient (DDPG) algorithms. Additionally the generalization ability of model under the dynamic disturbances and diverse terrains through domain randomization training and behavior cloning + staged reinforcement learning methods is studied. The simulation experiment results show that the proposed framework outperforms other algorithms in training efficiency, with the gait stability index( i.e., the standard deviation of the five-step center of mass height) of approximately 0.004 m. The robot can achieve the stable movement on complex terrains such as slopes and steps, whose motion performance is superior to that of the trained model. Finally, a preliminary prototype experiment was conducted to verify the feasibility and stability of proposed method in a real physical environment. In conclusion, the proposed deep reinforcement learning framework combing the biological gait imitation, behavior cloning initialization and phased reinforcement learning provides a new technical approach for the motion control of octopod robots.

    参考文献
    相似文献
    引证文献
引用本文

王诗雯,张军,温炜涛,李宏洋,宋爱国.基于深度强化学习的仿蟹机器人运动控制研究[J].仪器仪表学报,2026,47(5):71-85

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:
  • 最后修改日期:
  • 录用日期:
  • 在线发布日期: 2026-07-24
  • 出版日期:
文章二维码