Paper Reading AI Learner

The Monado SLAM Dataset for Egocentric Visual-Inertial Tracking

2025-07-31 18:28:07
Mateo de Mayo, Daniel Cremers, Taih\'u Pire

Abstract

Humanoid robots and mixed reality headsets benefit from the use of head-mounted sensors for tracking. While advancements in visual-inertial odometry (VIO) and simultaneous localization and mapping (SLAM) have produced new and high-quality state-of-the-art tracking systems, we show that these are still unable to gracefully handle many of the challenging settings presented in the head-mounted use cases. Common scenarios like high-intensity motions, dynamic occlusions, long tracking sessions, low-textured areas, adverse lighting conditions, saturation of sensors, to name a few, continue to be covered poorly by existing datasets in the literature. In this way, systems may inadvertently overlook these essential real-world issues. To address this, we present the Monado SLAM dataset, a set of real sequences taken from multiple virtual reality headsets. We release the dataset under a permissive CC BY 4.0 license, to drive advancements in VIO/SLAM research and development.

Abstract (translated)

人形机器人和混合现实头戴设备从使用头部安装传感器进行追踪中受益匪浅。尽管视觉惯性里程计(VIO)和即时定位与地图构建(SLAM)技术的进步产生了新的高质量最先进的追踪系统,但我们发现这些系统仍然无法优雅地应对许多出现在头部佩戴式应用中的挑战场景。常见的场景如高强度动作、动态遮挡、长时间的跟踪会话、低纹理区域、恶劣光照条件以及传感器饱和等问题,在现有文献中的数据集中仍被覆盖不足。因此,可能无意中忽略了这些问题在现实世界中的重要性。 为了解决这一问题,我们提出了Monado SLAM数据集,这是一个从多个虚拟现实头戴设备采集的真实序列集合。我们将该数据集以宽松的CC BY 4.0许可协议发布,旨在推动VIO/SLAM研究和开发领域的进步。

URL

https://arxiv.org/abs/2508.00088

PDF

https://arxiv.org/pdf/2508.00088.pdf


Tags
3D Action Action_Localization Action_Recognition Activity Adversarial Agent Attention Autonomous Bert Boundary_Detection Caption Chat Classification CNN Compressive_Sensing Contour Contrastive_Learning Deep_Learning Denoising Detection Dialog Diffusion Drone Dynamic_Memory_Network Edge_Detection Embedding Embodied Emotion Enhancement Face Face_Detection Face_Recognition Facial_Landmark Few-Shot Gait_Recognition GAN Gaze_Estimation Gesture Gradient_Descent Handwriting Human_Parsing Image_Caption Image_Classification Image_Compression Image_Enhancement Image_Generation Image_Matting Image_Retrieval Inference Inpainting Intelligent_Chip Knowledge Knowledge_Graph Language_Model LLM Matching Medical Memory_Networks Multi_Modal Multi_Task NAS NMT Object_Detection Object_Tracking OCR Ontology Optical_Character Optical_Flow Optimization Person_Re-identification Point_Cloud Portrait_Generation Pose Pose_Estimation Prediction QA Quantitative Quantitative_Finance Quantization Re-identification Recognition Recommendation Reconstruction Regularization Reinforcement_Learning Relation Relation_Extraction Represenation Represenation_Learning Restoration Review RNN Robot Salient Scene_Classification Scene_Generation Scene_Parsing Scene_Text Segmentation Self-Supervised Semantic_Instance_Segmentation Semantic_Segmentation Semi_Global Semi_Supervised Sence_graph Sentiment Sentiment_Classification Sketch SLAM Sparse Speech Speech_Recognition Style_Transfer Summarization Super_Resolution Surveillance Survey Text_Classification Text_Generation Time_Series Tracking Transfer_Learning Transformer Unsupervised Video_Caption Video_Classification Video_Indexing Video_Prediction Video_Retrieval Visual_Relation VQA Weakly_Supervised Zero-Shot