Paper Reading AI Learner

Vision-based 3D occupancy prediction in autonomous driving: a review and outlook

2024-05-04 07:39:25
Yanan Zhang, Jinqing Zhang, Zengran Wang, Junhao Xu, Di Huang

Abstract

In recent years, autonomous driving has garnered escalating attention for its potential to relieve drivers' burdens and improve driving safety. Vision-based 3D occupancy prediction, which predicts the spatial occupancy status and semantics of 3D voxel grids around the autonomous vehicle from image inputs, is an emerging perception task suitable for cost-effective perception system of autonomous driving. Although numerous studies have demonstrated the greater advantages of 3D occupancy prediction over object-centric perception tasks, there is still a lack of a dedicated review focusing on this rapidly developing field. In this paper, we first introduce the background of vision-based 3D occupancy prediction and discuss the challenges in this task. Secondly, we conduct a comprehensive survey of the progress in vision-based 3D occupancy prediction from three aspects: feature enhancement, deployment friendliness and label efficiency, and provide an in-depth analysis of the potentials and challenges of each category of methods. Finally, we present a summary of prevailing research trends and propose some inspiring future outlooks. To provide a valuable reference for researchers, a regularly updated collection of related papers, datasets, and codes is organized at this https URL.

Abstract (translated)

近年来,自动驾驶因为其减轻驾驶员负担和提高驾驶安全性的潜在优势而备受关注。基于视觉的3D占用预测,预测自动驾驶车辆周围3D体素网格的空间占用状态和语义,是一个适合自动驾驶低成本感知系统的 emerging perception 任务。尽管大量研究表明,与物体中心感知任务相比,3D占用预测具有更大的优势,但目前仍缺乏针对这一快速发展的领域的专门 review。在本文中,我们首先介绍了基于视觉的3D占用预测的背景,并讨论了这项任务的挑战。然后,我们从三个方面对基于视觉的3D占用预测的研究进展进行全面调查:特征增强、部署友好性和标签效率,并深入分析每种方法的潜在和挑战。最后,我们总结了当前研究趋势,并提出了鼓舞人心的未来展望。为了为研究人员提供有价值的参考,在 https://www. this URL 处组织了一个定期更新的相关论文、数据和代码的集合。

URL

https://arxiv.org/abs/2405.02595

PDF

https://arxiv.org/pdf/2405.02595.pdf


Tags
3D Action Action_Localization Action_Recognition Activity Adversarial Agent Attention Autonomous Bert Boundary_Detection Caption Chat Classification CNN Compressive_Sensing Contour Contrastive_Learning Deep_Learning Denoising Detection Dialog Diffusion Drone Dynamic_Memory_Network Edge_Detection Embedding Embodied Emotion Enhancement Face Face_Detection Face_Recognition Facial_Landmark Few-Shot Gait_Recognition GAN Gaze_Estimation Gesture Gradient_Descent Handwriting Human_Parsing Image_Caption Image_Classification Image_Compression Image_Enhancement Image_Generation Image_Matting Image_Retrieval Inference Inpainting Intelligent_Chip Knowledge Knowledge_Graph Language_Model LLM Matching Medical Memory_Networks Multi_Modal Multi_Task NAS NMT Object_Detection Object_Tracking OCR Ontology Optical_Character Optical_Flow Optimization Person_Re-identification Point_Cloud Portrait_Generation Pose Pose_Estimation Prediction QA Quantitative Quantitative_Finance Quantization Re-identification Recognition Recommendation Reconstruction Regularization Reinforcement_Learning Relation Relation_Extraction Represenation Represenation_Learning Restoration Review RNN Robot Salient Scene_Classification Scene_Generation Scene_Parsing Scene_Text Segmentation Self-Supervised Semantic_Instance_Segmentation Semantic_Segmentation Semi_Global Semi_Supervised Sence_graph Sentiment Sentiment_Classification Sketch SLAM Sparse Speech Speech_Recognition Style_Transfer Summarization Super_Resolution Surveillance Survey Text_Classification Text_Generation Tracking Transfer_Learning Transformer Unsupervised Video_Caption Video_Classification Video_Indexing Video_Prediction Video_Retrieval Visual_Relation VQA Weakly_Supervised Zero-Shot