Continuous Scene Representations for Embodied AI

2022-03-31 17:55:33

Samir Yitzhak Gadre, Kiana Ehsani, Shuran Song, Roozbeh Mottaghi

arXiv_CV

arXiv_CV Embedding Salient Relation Pose Embodied Agent

Abstract
Abstract (translated)
URL
PDF

Abstract

We propose Continuous Scene Representations (CSR), a scene representation constructed by an embodied agent navigating within a space, where objects and their relationships are modeled by continuous valued embeddings. Our method captures feature relationships between objects, composes them into a graph structure on-the-fly, and situates an embodied agent within the representation. Our key insight is to embed pair-wise relationships between objects in a latent space. This allows for a richer representation compared to discrete relations (e.g., [support], [next-to]) commonly used for building scene representations. CSR can track objects as the agent moves in a scene, update the representation accordingly, and detect changes in room configurations. Using CSR, we outperform state-of-the-art approaches for the challenging downstream task of visual room rearrangement, without any task specific training. Moreover, we show the learned embeddings capture salient spatial details of the scene and show applicability to real world data. A summery video and code is available at this https URL.

Abstract (translated)

URL

https://arxiv.org/abs/2203.17251

PDF

https://arxiv.org/pdf/2203.17251.pdf