Hierarchical World Model-Driven Visual Navigation Method for Embodied Robots in Dynamic Environments

Authors

  • Yujian Fu Qixin Honors School, Zhejiang Sci-Tech University, Hangzhou, China Author

DOI:

https://doi.org/10.70088/0b18tr67

Keywords:

embodied robot, visual navigation, world model, dynamic environment, risk prediction

Abstract

Visual navigation is a crucial capability for embodied robots to achieve autonomous movement in indoor environments. Due to factors such as moving obstacles, temporary obstructions, and changes in object positions, robots are prone to being disturbed during navigation, which affects driving safety and path efficiency. Traditional visual navigation methods are usually based on end-to-end strategies, semantic maps, or single-level world models. These approaches often face difficulties in balancing long-term navigation goals and short-term risk responses in dynamic environments. To address this issue, this paper proposes a visual navigation method based on a hierarchical world model, which models global semantic planning and local dynamic prediction separately. The global semantic world model is responsible for maintaining the semantic topological memory of the environment and providing support for sub-objective selection. The local dynamic world model predicts short-term reachability and potential contact risks between the robot and surrounding obstacles. Subsequently, the risk-aware hierarchical strategy integrates the value of sub-objectives, predicted reachability, and dynamic risks to generate navigation actions. Experiments were conducted on two simulation platforms, iGibson and Habitat-Matterport 3D (HM3D), covering static scenes and simulated dynamic scenes. The experimental results show that the proposed method outperforms existing methods in both static and dynamic scenarios, demonstrating better comprehensive performance in multiple aspects and verifying the effectiveness of the hierarchical world model in dynamic visual navigation tasks. This indicates that modeling semantic memory and dynamic risk prediction separately can help improve the navigation performance of robots in dynamic indoor environments.

Downloads

Published

2026-08-01