Paper Reading AI Learner

Motion is the Choreographer: Learning Latent Pose Dynamics for Seamless Sign Language Generation

2025-08-06 03:23:10
Jiayi He, Xu Wang, Shengeng Tang, Yaxiong Wang, Lechao Cheng, Dan Guo

Abstract

Sign language video generation requires producing natural signing motions with realistic appearances under precise semantic control, yet faces two critical challenges: excessive signer-specific data requirements and poor generalization. We propose a new paradigm for sign language video generation that decouples motion semantics from signer identity through a two-phase synthesis framework. First, we construct a signer-independent multimodal motion lexicon, where each gloss is stored as identity-agnostic pose, gesture, and 3D mesh sequences, requiring only one recording per sign. This compact representation enables our second key innovation: a discrete-to-continuous motion synthesis stage that transforms retrieved gloss sequences into temporally coherent motion trajectories, followed by identity-aware neural rendering to produce photorealistic videos of arbitrary signers. Unlike prior work constrained by signer-specific datasets, our method treats motion as a first-class citizen: the learned latent pose dynamics serve as a portable "choreography layer" that can be visually realized through different human appearances. Extensive experiments demonstrate that disentangling motion from identity is not just viable but advantageous - enabling both high-quality synthesis and unprecedented flexibility in signer personalization.

Abstract (translated)

手语视频生成需要在精确的语义控制下产生自然的手势动作和逼真的外观,但面临着两个关键挑战:对手势特定数据的需求过多以及泛化能力差。我们提出了一种新的手语视频生成范式,通过一个两阶段合成框架将手势语义与执行者身份解耦。 首先,我们构建了一个与执行者无关的多模态动作词典,在其中每个词(gloss)都以无身份标识的姿态、手势和3D网格序列的形式存储。这种表示方式只需要为每个手语记录一次数据即可。这使得我们的第二个关键创新成为可能:从检索到的手势序列中进行离散到连续的动作合成阶段,随后通过面向执行者外观的神经渲染生成任意执行者的逼真视频。 不同于之前受限于特定执行者数据集的方法,我们的方法将动作视为重要组成部分:“学习到的姿态动力学”可以作为便携式的“编舞层”,可以在不同的人类外表中视觉实现。广泛的实验表明,分离手势和身份不仅可行而且具有优势——这使得高质量的合成与前所未有的个性化成为可能。

URL

https://arxiv.org/abs/2508.04049

PDF

https://arxiv.org/pdf/2508.04049.pdf


Tags
3D Action Action_Localization Action_Recognition Activity Adversarial Agent Attention Autonomous Bert Boundary_Detection Caption Chat Classification CNN Compressive_Sensing Contour Contrastive_Learning Deep_Learning Denoising Detection Dialog Diffusion Drone Dynamic_Memory_Network Edge_Detection Embedding Embodied Emotion Enhancement Face Face_Detection Face_Recognition Facial_Landmark Few-Shot Gait_Recognition GAN Gaze_Estimation Gesture Gradient_Descent Handwriting Human_Parsing Image_Caption Image_Classification Image_Compression Image_Enhancement Image_Generation Image_Matting Image_Retrieval Inference Inpainting Intelligent_Chip Knowledge Knowledge_Graph Language_Model LLM Matching Medical Memory_Networks Multi_Modal Multi_Task NAS NMT Object_Detection Object_Tracking OCR Ontology Optical_Character Optical_Flow Optimization Person_Re-identification Point_Cloud Portrait_Generation Pose Pose_Estimation Prediction QA Quantitative Quantitative_Finance Quantization Re-identification Recognition Recommendation Reconstruction Regularization Reinforcement_Learning Relation Relation_Extraction Represenation Represenation_Learning Restoration Review RNN Robot Salient Scene_Classification Scene_Generation Scene_Parsing Scene_Text Segmentation Self-Supervised Semantic_Instance_Segmentation Semantic_Segmentation Semi_Global Semi_Supervised Sence_graph Sentiment Sentiment_Classification Sketch SLAM Sparse Speech Speech_Recognition Style_Transfer Summarization Super_Resolution Surveillance Survey Text_Classification Text_Generation Time_Series Tracking Transfer_Learning Transformer Unsupervised Video_Caption Video_Classification Video_Indexing Video_Prediction Video_Retrieval Visual_Relation VQA Weakly_Supervised Zero-Shot