Abstract
Sign language video generation requires producing natural signing motions with realistic appearances under precise semantic control, yet faces two critical challenges: excessive signer-specific data requirements and poor generalization. We propose a new paradigm for sign language video generation that decouples motion semantics from signer identity through a two-phase synthesis framework. First, we construct a signer-independent multimodal motion lexicon, where each gloss is stored as identity-agnostic pose, gesture, and 3D mesh sequences, requiring only one recording per sign. This compact representation enables our second key innovation: a discrete-to-continuous motion synthesis stage that transforms retrieved gloss sequences into temporally coherent motion trajectories, followed by identity-aware neural rendering to produce photorealistic videos of arbitrary signers. Unlike prior work constrained by signer-specific datasets, our method treats motion as a first-class citizen: the learned latent pose dynamics serve as a portable "choreography layer" that can be visually realized through different human appearances. Extensive experiments demonstrate that disentangling motion from identity is not just viable but advantageous - enabling both high-quality synthesis and unprecedented flexibility in signer personalization.
Abstract (translated)
手语视频生成需要在精确的语义控制下产生自然的手势动作和逼真的外观,但面临着两个关键挑战:对手势特定数据的需求过多以及泛化能力差。我们提出了一种新的手语视频生成范式,通过一个两阶段合成框架将手势语义与执行者身份解耦。 首先,我们构建了一个与执行者无关的多模态动作词典,在其中每个词(gloss)都以无身份标识的姿态、手势和3D网格序列的形式存储。这种表示方式只需要为每个手语记录一次数据即可。这使得我们的第二个关键创新成为可能:从检索到的手势序列中进行离散到连续的动作合成阶段,随后通过面向执行者外观的神经渲染生成任意执行者的逼真视频。 不同于之前受限于特定执行者数据集的方法,我们的方法将动作视为重要组成部分:“学习到的姿态动力学”可以作为便携式的“编舞层”,可以在不同的人类外表中视觉实现。广泛的实验表明,分离手势和身份不仅可行而且具有优势——这使得高质量的合成与前所未有的个性化成为可能。
URL
https://arxiv.org/abs/2508.04049