Abstract
Sign language, which contains hand movements, facial expressions and bodily gestures, is a significant medium for communicating with hard-of-hearing people. A well-trained sign language community communicates easily, but those who don't know sign language face significant challenges. Recognition and generation are basic communication methods between hearing and hard-of-hearing individuals. Despite progress in recognition, sign language generation still needs to be explored. The Progressive Growing of Generative Adversarial Network (ProGAN) excels at producing high-quality images, while the Self-Attention Generative Adversarial Network (SAGAN) generates feature-rich images at medium resolutions. Balancing resolution and detail is crucial for sign language image generation. We are developing a Generative Adversarial Network (GAN) variant that combines both models to generate feature-rich, high-resolution, and class-conditional sign language images. Our modified Attention-based model generates high-quality images of Indian Sign Language letters, numbers, and words, outperforming the traditional ProGAN in Inception Score (IS) and Fréchet Inception Distance (FID), with improvements of 3.2 and 30.12, respectively. Additionally, we are publishing a large dataset incorporating high-quality images of Indian Sign Language alphabets, numbers, and 129 words.
Abstract (translated)
手语包含手势、面部表情和身体动作,是与听力障碍人士沟通的重要媒介。经过良好训练的手语社区能够轻松交流,但对于不懂手语的人来说则面临巨大挑战。识别和生成是听觉健全者和听力障碍者之间基本的交流方式。尽管在识别方面取得了进展,但手语生成仍需进一步探索。 渐进式生长生成对抗网络(ProGAN)擅长生成高质量图像,而自注意力生成对抗网络(SAGAN)则擅长在中等分辨率下生成细节丰富的图像。对于手语图像生成而言,在平衡分辨率和细节度上至关重要。我们正在开发一种结合这两种模型的生成对抗网络(GAN)变体,以生成特征丰富、高分辨率且具有类别条件的手语图像。 我们的基于注意力机制的改良模型能够生成高质量的印度手语字母、数字及词汇图像,在Inception分数(IS)和Fréchet Inception距离(FID)指标上超越了传统的ProGAN,分别提高了3.2和30.12。此外,我们还计划发布一个包含高分辨率印度手语字母、数字以及129个词语的大规模数据集。
URL
https://arxiv.org/abs/2508.09522