Paper Reading AI Learner

CoCoCo: Improving Text-Guided Video Inpainting for Better Consistency, Controllability and Compatibility

2024-03-18 17:59:27
Bojia Zi, Shihao Zhao, Xianbiao Qi, Jianan Wang, Yukai Shi, Qianyu Chen, Bin Liang, Kam-Fai Wong, Lei Zhang

Abstract

Recent advancements in video generation have been remarkable, yet many existing methods struggle with issues of consistency and poor text-video alignment. Moreover, the field lacks effective techniques for text-guided video inpainting, a stark contrast to the well-explored domain of text-guided image inpainting. To this end, this paper proposes a novel text-guided video inpainting model that achieves better consistency, controllability and compatibility. Specifically, we introduce a simple but efficient motion capture module to preserve motion consistency, and design an instance-aware region selection instead of a random region selection to obtain better textual controllability, and utilize a novel strategy to inject some personalized models into our CoCoCo model and thus obtain better model compatibility. Extensive experiments show that our model can generate high-quality video clips. Meanwhile, our model shows better motion consistency, textual controllability and model compatibility. More details are shown in [this http URL](this http URL).

Abstract (translated)

近年来在视频生成方面的进展令人印象深刻,然而,许多现有方法在一致性和文本-视频对齐方面存在问题。此外,该领域缺乏有效的文本指导视频修复技术,与文本指导图像修复领域已被充分探索的领域形成鲜明对比。为此,本文提出了一种新颖的文本指导视频修复模型,实现了更好的一致性、可控制性和兼容性。具体来说,我们引入了一个简单但高效的动作捕捉模块来保留运动一致性,并设计了一个实例感知区域选择,而不是随机区域选择,以获得更好的文本控制性,并利用一种新颖的方法将一些个性化的模型注入到我们的CoCoCo模型中,从而实现更好的模型兼容性。大量实验结果表明,我们的模型可以生成高质量的视频剪辑。同时,我们的模型在运动一致性、文本控制性和模型兼容性方面表现更好。更多细节可见于[http://www.thisurl.com](http://www.thisurl.com)。

URL

https://arxiv.org/abs/2403.12035

PDF

https://arxiv.org/pdf/2403.12035.pdf


Tags
3D Action Action_Localization Action_Recognition Activity Adversarial Agent Attention Autonomous Bert Boundary_Detection Caption Chat Classification CNN Compressive_Sensing Contour Contrastive_Learning Deep_Learning Denoising Detection Dialog Diffusion Drone Dynamic_Memory_Network Edge_Detection Embedding Embodied Emotion Enhancement Face Face_Detection Face_Recognition Facial_Landmark Few-Shot Gait_Recognition GAN Gaze_Estimation Gesture Gradient_Descent Handwriting Human_Parsing Image_Caption Image_Classification Image_Compression Image_Enhancement Image_Generation Image_Matting Image_Retrieval Inference Inpainting Intelligent_Chip Knowledge Knowledge_Graph Language_Model LLM Matching Medical Memory_Networks Multi_Modal Multi_Task NAS NMT Object_Detection Object_Tracking OCR Ontology Optical_Character Optical_Flow Optimization Person_Re-identification Point_Cloud Portrait_Generation Pose Pose_Estimation Prediction QA Quantitative Quantitative_Finance Quantization Re-identification Recognition Recommendation Reconstruction Regularization Reinforcement_Learning Relation Relation_Extraction Represenation Represenation_Learning Restoration Review RNN Robot Salient Scene_Classification Scene_Generation Scene_Parsing Scene_Text Segmentation Self-Supervised Semantic_Instance_Segmentation Semantic_Segmentation Semi_Global Semi_Supervised Sence_graph Sentiment Sentiment_Classification Sketch SLAM Sparse Speech Speech_Recognition Style_Transfer Summarization Super_Resolution Surveillance Survey Text_Classification Text_Generation Tracking Transfer_Learning Transformer Unsupervised Video_Caption Video_Classification Video_Indexing Video_Prediction Video_Retrieval Visual_Relation VQA Weakly_Supervised Zero-Shot