Abstract
High-precision scene parsing tasks, including image matting and dichotomous segmentation, aim to accurately predict masks with extremely fine details (such as hair). Most existing methods focus on salient, single foreground objects. While interactive methods allow for target adjustment, their class-agnostic design restricts generalization across different categories. Furthermore, the scarcity of high-quality annotation has led to a reliance on inharmonious synthetic data, resulting in poor generalization to real-world scenarios. To this end, we propose a Foreground Consistent Learning model, dubbed as FCLM, to address the aforementioned issues. Specifically, we first introduce a Depth-Aware Distillation strategy where we transfer the depth-related knowledge for better foreground representation. Considering the data dilemma, we term the processing of synthetic data as domain adaptation problem where we propose a domain-invariant learning strategy to focus on foreground learning. To support interactive prediction, we contribute an Object-Oriented Decoder that can receive both visual and language prompts to predict the referring target. Experimental results show that our method quantitatively and qualitatively outperforms SOTA methods.
Abstract (translated)
高精度场景解析任务,包括图像抠图和二元分割,旨在准确预测包含极细微节(如头发)的掩码。现有大多数方法主要关注显著、单一前景对象。尽管交互式方法允许目标调整,但它们无类别感知的设计限制了跨不同类别的泛化能力。此外,高质量标注数据的稀缺导致对不协调合成数据的高度依赖,从而在真实场景中的表现不佳。 为了解决上述问题,我们提出了一种名为FCLM(Foreground Consistent Learning model)的方法来解决这些问题。具体来说,我们首先引入了深度感知蒸馏策略,在该策略中我们转移与深度相关的知识以更好地表示前景对象。考虑到数据困境,我们将合成数据的处理视为领域适应问题,并提出了域不变学习策略专注于前景学习。为了支持交互式预测,我们贡献了一个面向对象解码器,它可以接收视觉和语言提示来预测引用目标。 实验结果表明,我们的方法在定量和定性上均优于现有最先进的方法(SOTA)。
URL
https://arxiv.org/abs/2601.12080