Abstract
Object navigation in open-world environments remains a formidable and pervasive challenge for robotic systems, particularly when it comes to executing long-horizon tasks that require both open-world object detection and high-level task planning. Traditional methods often struggle to integrate these components effectively, and this limits their capability to deal with complex, long-range navigation missions. In this paper, we propose LOVON, a novel framework that integrates large language models (LLMs) for hierarchical task planning with open-vocabulary visual detection models, tailored for effective long-range object navigation in dynamic, unstructured environments. To tackle real-world challenges including visual jittering, blind zones, and temporary target loss, we design dedicated solutions such as Laplacian Variance Filtering for visual stabilization. We also develop a functional execution logic for the robot that guarantees LOVON's capabilities in autonomous navigation, task adaptation, and robust task completion. Extensive evaluations demonstrate the successful completion of long-sequence tasks involving real-time detection, search, and navigation toward open-vocabulary dynamic targets. Furthermore, real-world experiments across different legged robots (Unitree Go2, B2, and H1-2) showcase the compatibility and appealing plug-and-play feature of LOVON.
Abstract (translated)
在开放世界环境中,对象导航对机器人系统来说仍然是一个艰巨且普遍的挑战,特别是在执行需要开放式世界目标检测和高层次任务规划的长时任务时。传统方法往往难以有效整合这些组件,从而限制了它们处理复杂、远程导航任务的能力。本文提出了一种新的框架LOVON,该框架结合了大型语言模型(LLMs)进行分层任务规划与开放词汇视觉检测模型,并针对动态和无结构环境中的长期目标导航进行了优化设计。 为了应对包括视觉抖动、盲区和临时目标丢失在内的现实世界挑战,我们专门设计了解决方案如拉普拉斯方差过滤器以实现图像稳定。此外,还开发了一种功能执行逻辑,确保LOVON在自主导航、任务适应性和稳健的任务完成方面的能力。广泛的评估表明,在实时检测、搜索和导航向开放词汇动态目标的长序列任务中能够成功完成这些任务。 跨不同腿式机器人(如Unitree Go2、B2和H1-2)的真实世界实验展示了LOVON的高度兼容性及其引人注目的即插即用特性。
URL
https://arxiv.org/abs/2507.06747