Paper Reading AI Learner

AID4AD: Aerial Image Data for Automated Driving Perception

2025-08-04 07:38:18
Daniel Lengerer, Mathias Pechinger, Klaus Bogenberger, Carsten Markgraf

Abstract

This work investigates the integration of spatially aligned aerial imagery into perception tasks for automated vehicles (AVs). As a central contribution, we present AID4AD, a publicly available dataset that augments the nuScenes dataset with high-resolution aerial imagery precisely aligned to its local coordinate system. The alignment is performed using SLAM-based point cloud maps provided by nuScenes, establishing a direct link between aerial data and nuScenes local coordinate system. To ensure spatial fidelity, we propose an alignment workflow that corrects for localization and projection distortions. A manual quality control process further refines the dataset by identifying a set of high-quality alignments, which we publish as ground truth to support future research on automated registration. We demonstrate the practical value of AID4AD in two representative tasks: in online map construction, aerial imagery serves as a complementary input that improves the mapping process; in motion prediction, it functions as a structured environmental representation that replaces high-definition maps. Experiments show that aerial imagery leads to a 15-23% improvement in map construction accuracy and a 2% gain in trajectory prediction performance. These results highlight the potential of aerial imagery as a scalable and adaptable source of environmental context in automated vehicle systems, particularly in scenarios where high-definition maps are unavailable, outdated, or costly to maintain. AID4AD, along with evaluation code and pretrained models, is publicly released to foster further research in this direction: this https URL.

Abstract (translated)

这项工作探讨了将空间对齐的航拍图像集成到自动驾驶汽车(AV)感知任务中的方法。作为核心贡献,我们提出了AID4AD,这是一个公开可用的数据集,它通过高分辨率的、与nuScenes数据集中本地坐标系统精确对齐的航拍图像来扩充nuScenes数据集。该对准过程使用了由nuScenes提供的基于SLAM的点云地图,从而在空中数据和nuScenes本地坐标系之间建立了直接链接。 为了确保空间精度,我们提出了一种对准工作流程,用于纠正定位和投影失真。通过手动质量控制过程进一步完善数据集,识别出一系列高质量对准,并将其发布为地面实况以支持未来有关自动注册的研究。我们在两个代表性任务中展示了AID4AD的实际价值:在线地图构建过程中,航拍图像作为补充输入改进了制图流程;在运动预测方面,它充当结构化的环境表示来替代高精度地图。 实验表明,使用航拍图像可以将地图构造准确度提高15-23%,并将轨迹预测性能提升2%。这些结果突显了航空影像作为一种可扩展且适应性强的环境上下文数据源,在自动驾驶系统中的潜在价值,特别是在没有、过时或维护成本高昂的高精度地图的情况下。 AID4AD以及评估代码和预训练模型已公开发布,以促进该领域的进一步研究:[此链接](this https URL)。

URL

https://arxiv.org/abs/2508.02140

PDF

https://arxiv.org/pdf/2508.02140.pdf


Tags
3D Action Action_Localization Action_Recognition Activity Adversarial Agent Attention Autonomous Bert Boundary_Detection Caption Chat Classification CNN Compressive_Sensing Contour Contrastive_Learning Deep_Learning Denoising Detection Dialog Diffusion Drone Dynamic_Memory_Network Edge_Detection Embedding Embodied Emotion Enhancement Face Face_Detection Face_Recognition Facial_Landmark Few-Shot Gait_Recognition GAN Gaze_Estimation Gesture Gradient_Descent Handwriting Human_Parsing Image_Caption Image_Classification Image_Compression Image_Enhancement Image_Generation Image_Matting Image_Retrieval Inference Inpainting Intelligent_Chip Knowledge Knowledge_Graph Language_Model LLM Matching Medical Memory_Networks Multi_Modal Multi_Task NAS NMT Object_Detection Object_Tracking OCR Ontology Optical_Character Optical_Flow Optimization Person_Re-identification Point_Cloud Portrait_Generation Pose Pose_Estimation Prediction QA Quantitative Quantitative_Finance Quantization Re-identification Recognition Recommendation Reconstruction Regularization Reinforcement_Learning Relation Relation_Extraction Represenation Represenation_Learning Restoration Review RNN Robot Salient Scene_Classification Scene_Generation Scene_Parsing Scene_Text Segmentation Self-Supervised Semantic_Instance_Segmentation Semantic_Segmentation Semi_Global Semi_Supervised Sence_graph Sentiment Sentiment_Classification Sketch SLAM Sparse Speech Speech_Recognition Style_Transfer Summarization Super_Resolution Surveillance Survey Text_Classification Text_Generation Time_Series Tracking Transfer_Learning Transformer Unsupervised Video_Caption Video_Classification Video_Indexing Video_Prediction Video_Retrieval Visual_Relation VQA Weakly_Supervised Zero-Shot