基于三维几何世界模型的具身智能方法研究
DOI:
CSTR:
作者:
作者单位:

1.东北大学信息科学与工程学院沈阳110169; 2.东北大学机器人科学与工程学院沈阳110169

作者简介:

通讯作者:

中图分类号:

TH86TP242.6

基金项目:

工业机器人用具身智能大模型平台研究与示范应用(2024QK2001)项目资助


Research on embodied AI methods based on 3D geometric world models
Author:
Affiliation:

1.School of Information Science and Engineering, Northeastern University, Shenyang 110169, China; 2.School of Robot Science and Engineering, Northeastern University, Shenyang 110169, China

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    现有的具身智能操控方法大多采用端到端的学习策略,虽在特定任务上表现尚可,但因将物理交互视为黑箱问题,牺牲了系统可解释性,且面临几何感知缺失、长程规划累积误差大及动力学适应性差等挑战。为解决非结构化场景下的精细操作难题,提出了一种基于三维几何世界模型(3D-GWM)与条件流匹配策略的具身智能操控框架,包含三维感知与压缩、生成式几何世界模型和条件流匹配动作生成3个核心模块。感知模块提出了互补双域特征增强算法,综合提取空间高频细节与全局频域上下文,并结合指令驱动的动态词元压缩机制,在保证高保真特征提取的同时降低计算开销。部署评估中,输入视觉词元数由2 048压缩至128,峰值显存占用由58.7 GB降至17.2 GB。认知模块采用中间帧预测策略,通过预测子任务阶段性的三维几何终态,为长程任务提供具备因果逻辑的中间目标引导,有效抑制规划过程中的累积误差。执行模块采用条件流匹配策略,利用几何锚点构建目标向量场,生成符合物理接触约束的平滑动力学轨迹,降低刚性逆运动学求解中奇异性问题的发生风险,并缓解轨迹不连续带来的执行不稳定。实验结果表明,3D-GWM在复杂逻辑与高精度任务中的仿真平均成功率为89.4%,真机平均成功率为82.0%;相较于当前代表性基线方法,平均成功率分别提升7.0%和8.75%。

    Abstract:

    Most existing embodied manipulation methods predominantly rely on end-to-end learning strategies. Although performing adequately on specific tasks, these approaches often treat physical interaction as a black box, thereby sacrificing interpretability and suffering from insufficient geometric grounding, large error accumulation in long-horizon planning, and weak adaptability to complex dynamics. To address fine-grained manipulation in unstructured scenarios, this article proposes an embodied manipulation framework based on a 3D geometric world model (3D-GWM) and conditional flow matching, termed 3D-GWM, which comprises three modules: 3D perception and compression, a generative geometric world model, and conditional flow-matching-based action generation. In the perception module, a complementary dual-domain feature enhancement algorithm is proposed to jointly capture spatial high-frequency details and global frequency-domain context, together with an instruction-driven dynamic token compression mechanism to reduce computation while preserving feature fidelity. In deployment evaluations, the number of input visual tokens is reduced from 2 048 to 128, while the peak GPU memory usage is reduced from 58.7 GB to 17.2 GB. In the cognition module, an intermediate-frame prediction strategy is adopted. By predicting stage-wise 3D geometric terminal states of sub-tasks, it provides intermediate goal guidance with causal logic for long-horizon manipulation and mitigates error accumulation during planning. In the execution module, a conditional flow matching strategy is employed. By using geometric anchors to construct target vector fields, it generates smooth trajectories that satisfy physical contact constraints and reduces the risk of singularities commonly encountered in rigid inverse-kinematics solvers. Experimental results show that 3D-GWM achieves an average success rate of 89.4% in simulation and 82.0% on the real robot for complex-logic and high-precision tasks; compared with representative baseline methods, the average success rate increases by 7.0% and 8.75%, respectively.

    参考文献
    相似文献
    引证文献
引用本文

覃宏伟,姜杨,白佳硕,徐天尧,杨轩溥.基于三维几何世界模型的具身智能方法研究[J].仪器仪表学报,2026,47(5):23-33

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:
  • 最后修改日期:
  • 录用日期:
  • 在线发布日期: 2026-07-24
  • 出版日期:
文章二维码