Paper Reading AI Learner

$S^2M^2$: Scalable Stereo Matching Model for Reliable Depth Estimation

2025-07-17 15:40:18
Junhong Min, Youngpil Jeon, Jimin Kim, Minyong Choi

Abstract

The pursuit of a generalizable stereo matching model, capable of performing across varying resolutions and disparity ranges without dataset-specific fine-tuning, has revealed a fundamental trade-off. Iterative local search methods achieve high scores on constrained benchmarks, but their core mechanism inherently limits the global consistency required for true generalization. On the other hand, global matching architectures, while theoretically more robust, have been historically rendered infeasible by prohibitive computational and memory costs. We resolve this dilemma with $S^2M^2$: a global matching architecture that achieves both state-of-the-art accuracy and high efficiency without relying on cost volume filtering or deep refinement stacks. Our design integrates a multi-resolution transformer for robust long-range correspondence, trained with a novel loss function that concentrates probability on feasible matches. This approach enables a more robust joint estimation of disparity, occlusion, and confidence. $S^2M^2$ establishes a new state of the art on the Middlebury v3 and ETH3D benchmarks, significantly outperforming prior methods across most metrics while reconstructing high-quality details with competitive efficiency.

Abstract (translated)

追求一种通用的立体匹配模型,这种模型能够在不同分辨率和视差范围内工作,并且不需要特定数据集的微调,已经揭示了一个基本的权衡。迭代局部搜索方法在受限制的基准测试中取得了高分,但其核心机制内在地限制了实现全局一致性的能力,这是真正泛化所必需的。另一方面,虽然理论上更稳健的全球匹配架构被认为是可行的,但由于计算和内存成本过高,在历史上一直难以实施。 我们通过提出$S^2M^2$解决了这一难题:这是一种全局匹配架构,能够在不依赖于代价体积过滤或深度细化堆栈的情况下实现最先进的精度和高效率。我们的设计集成了一个多分辨率变压器,用于稳健的长距离对应,并使用一种新的损失函数进行训练,该函数专注于可行匹配的概率集中。这种方法使视差、遮挡和置信度的一致估计更加稳健。 $S^2M^2$在Middlebury v3和ETH3D基准测试中建立了新的状态,相对于先前的方法,在大多数指标上表现出了显著的优越性,并且能够以具有竞争力的效率重建高质量的细节。

URL

https://arxiv.org/abs/2507.13229

PDF

https://arxiv.org/pdf/2507.13229.pdf


Tags
3D Action Action_Localization Action_Recognition Activity Adversarial Agent Attention Autonomous Bert Boundary_Detection Caption Chat Classification CNN Compressive_Sensing Contour Contrastive_Learning Deep_Learning Denoising Detection Dialog Diffusion Drone Dynamic_Memory_Network Edge_Detection Embedding Embodied Emotion Enhancement Face Face_Detection Face_Recognition Facial_Landmark Few-Shot Gait_Recognition GAN Gaze_Estimation Gesture Gradient_Descent Handwriting Human_Parsing Image_Caption Image_Classification Image_Compression Image_Enhancement Image_Generation Image_Matting Image_Retrieval Inference Inpainting Intelligent_Chip Knowledge Knowledge_Graph Language_Model LLM Matching Medical Memory_Networks Multi_Modal Multi_Task NAS NMT Object_Detection Object_Tracking OCR Ontology Optical_Character Optical_Flow Optimization Person_Re-identification Point_Cloud Portrait_Generation Pose Pose_Estimation Prediction QA Quantitative Quantitative_Finance Quantization Re-identification Recognition Recommendation Reconstruction Regularization Reinforcement_Learning Relation Relation_Extraction Represenation Represenation_Learning Restoration Review RNN Robot Salient Scene_Classification Scene_Generation Scene_Parsing Scene_Text Segmentation Self-Supervised Semantic_Instance_Segmentation Semantic_Segmentation Semi_Global Semi_Supervised Sence_graph Sentiment Sentiment_Classification Sketch SLAM Sparse Speech Speech_Recognition Style_Transfer Summarization Super_Resolution Surveillance Survey Text_Classification Text_Generation Time_Series Tracking Transfer_Learning Transformer Unsupervised Video_Caption Video_Classification Video_Indexing Video_Prediction Video_Retrieval Visual_Relation VQA Weakly_Supervised Zero-Shot