Abstract
We propose SLARM, a feed-forward model that unifies dynamic scene reconstruction, semantic understanding, and real-time streaming inference. SLARM captures complex, non-uniform motion through higher-order motion modeling, trained solely on differentiable renderings without any flow supervision. Besides, SLARM distills semantic features from LSeg to obtain language-aligned representations. This design enables semantic querying via natural language, and the tight coupling between semantics and geometry further enhances the accuracy and robustness of dynamic reconstruction. Moreover, SLARM processes image sequences using window-based causal attention, achieving stable, low-latency streaming inference without accumulating memory cost. Within this unified framework, SLARM achieves state-of-the-art results in dynamic estimation, rendering quality, and scene parsing, improving motion accuracy by 21%, reconstruction PSNR by 1.6 dB, and segmentation mIoU by 20% over existing methods.
Abstract (translated)
我们提出SLARM,一种统一动态场景重建、语义理解与实时流式推理的前馈模型。SLARM通过高阶运动建模捕捉复杂非均匀运动,仅使用可微渲染进行训练而无需光流监督。此外,SLARM从LSeg中提炼语义特征以获得语言对齐表征,该设计支持通过自然语言进行语义查询,且语义与几何的紧密耦合进一步提升了动态重建的精度与鲁棒性。同时,SLARM采用基于窗口的因果注意力处理图像序列,实现了稳定低延迟的流式推理且无需累积内存成本。在此统一框架下,SLARM在动态估计、渲染质量与场景解析任务中均达到最优性能,相较于现有方法,运动精度提升21%,重建PSNR提升1.6dB,分割mIoU提升20%。
URL
https://arxiv.org/abs/2603.22893