Paper Reading AI Learner

Bridging the Last Mile in Sim-to-Real Robot Perception via Bayesian Active Learning

2021-09-23 14:45:40
Jianxiang Feng, Jongseok Lee, Maximilian Durner, Rudolph Triebel

Abstract

Learning from synthetic data is popular in avariety of robotic vision tasks such as object detection, becauselarge amount of data can be generated without annotationsby humans. However, when relying only on synthetic data,we encounter the well-known problem of the simulation-to-reality (Sim-to-Real) gap, which is hard to resolve completelyin practice. For such cases, real human-annotated data isnecessary to bridge this gap, and in our work we focus on howto acquire this data efficiently. Therefore, we propose a Sim-to-Real pipeline that relies on deep Bayesian active learningand aims to minimize the manual annotation efforts. We devisea learning paradigm that autonomously selects the data thatis considered useful for the human expert to annotate. Toachieve this, a Bayesian Neural Network (BNN) object detectorproviding reliable uncertain estimates is adapted to infer theinformativeness of the unlabeled data, in order to performactive learning. In our experiments on two object detectiondata sets, we show that the labeling effort required to bridge thereality gap can be reduced to a small amount. Furthermore, wedemonstrate the practical effectiveness of this idea in a graspingtask on an assistive robot.

Abstract (translated)

URL

https://arxiv.org/abs/2109.11547

PDF

https://arxiv.org/pdf/2109.11547.pdf


Tags
3D Action Action_Localization Action_Recognition Activity Adversarial Agent Attention Autonomous Bert Boundary_Detection Caption Chat Classification CNN Compressive_Sensing Contour Contrastive_Learning Deep_Learning Denoising Detection Dialog Diffusion Drone Dynamic_Memory_Network Edge_Detection Embedding Embodied Emotion Enhancement Face Face_Detection Face_Recognition Facial_Landmark Few-Shot Gait_Recognition GAN Gaze_Estimation Gesture Gradient_Descent Handwriting Human_Parsing Image_Caption Image_Classification Image_Compression Image_Enhancement Image_Generation Image_Matting Image_Retrieval Inference Inpainting Intelligent_Chip Knowledge Knowledge_Graph Language_Model Matching Medical Memory_Networks Multi_Modal Multi_Task NAS NMT Object_Detection Object_Tracking OCR Ontology Optical_Character Optical_Flow Optimization Person_Re-identification Point_Cloud Portrait_Generation Pose Pose_Estimation Prediction QA Quantitative Quantitative_Finance Quantization Re-identification Recognition Recommendation Reconstruction Regularization Reinforcement_Learning Relation Relation_Extraction Represenation Represenation_Learning Restoration Review RNN Salient Scene_Classification Scene_Generation Scene_Parsing Scene_Text Segmentation Self-Supervised Semantic_Instance_Segmentation Semantic_Segmentation Semi_Global Semi_Supervised Sence_graph Sentiment Sentiment_Classification Sketch SLAM Sparse Speech Speech_Recognition Style_Transfer Summarization Super_Resolution Surveillance Survey Text_Classification Text_Generation Tracking Transfer_Learning Transformer Unsupervised Video_Caption Video_Classification Video_Indexing Video_Prediction Video_Retrieval Visual_Relation VQA Weakly_Supervised Zero-Shot