Abstract
Video Scene Parsing (VSP) has emerged as a cornerstone in computer vision, facilitating the simultaneous segmentation, recognition, and tracking of diverse visual entities in dynamic scenes. In this survey, we present a holistic review of recent advances in VSP, covering a wide array of vision tasks, including Video Semantic Segmentation (VSS), Video Instance Segmentation (VIS), Video Panoptic Segmentation (VPS), as well as Video Tracking and Segmentation (VTS), and Open-Vocabulary Video Segmentation (OVVS). We systematically analyze the evolution from traditional hand-crafted features to modern deep learning paradigms -- spanning from fully convolutional networks to the latest transformer-based architectures -- and assess their effectiveness in capturing both local and global temporal contexts. Furthermore, our review critically discusses the technical challenges, ranging from maintaining temporal consistency to handling complex scene dynamics, and offers a comprehensive comparative study of datasets and evaluation metrics that have shaped current benchmarking standards. By distilling the key contributions and shortcomings of state-of-the-art methodologies, this survey highlights emerging trends and prospective research directions that promise to further elevate the robustness and adaptability of VSP in real-world applications.
Abstract (translated)
视频场景解析(VSP)已成为计算机视觉领域的基石,它能够同时对动态场景中的各种视觉实体进行分割、识别和跟踪。本文综述了近期在VSP领域取得的进展,涵盖了包括视频语义分割(VSS)、视频实例分割(VIS)、视频全景分割(VPS),以及视频跟踪与分割(VTS)和开放词汇视频分割(OVVS)在内的广泛视觉任务。我们系统地分析了从传统的手工设计特征到现代深度学习范式的演变过程,涵盖了从全卷积网络到最新基于变换器的架构,并评估它们在捕捉局部和全局时间上下文方面的有效性。此外,我们的综述还批判性地讨论了技术挑战,包括保持时间一致性以及处理复杂的场景动态变化,并提供了对塑造当前基准标准的数据集和评价指标的全面比较研究。通过提炼现有最先进方法的关键贡献与不足之处,本文概述了正在形成的新趋势及具有前景的研究方向,这些方向有望进一步提升VSP在实际应用中的鲁棒性和适应性。
URL
https://arxiv.org/abs/2506.13552