Paper Reading AI Learner

Generation of Indian Sign Language Letters, Numbers, and Words

2025-08-13 06:10:20
Ajeet Kumar Yadav, Nishant Kumar, Rathna G N

Abstract

Sign language, which contains hand movements, facial expressions and bodily gestures, is a significant medium for communicating with hard-of-hearing people. A well-trained sign language community communicates easily, but those who don't know sign language face significant challenges. Recognition and generation are basic communication methods between hearing and hard-of-hearing individuals. Despite progress in recognition, sign language generation still needs to be explored. The Progressive Growing of Generative Adversarial Network (ProGAN) excels at producing high-quality images, while the Self-Attention Generative Adversarial Network (SAGAN) generates feature-rich images at medium resolutions. Balancing resolution and detail is crucial for sign language image generation. We are developing a Generative Adversarial Network (GAN) variant that combines both models to generate feature-rich, high-resolution, and class-conditional sign language images. Our modified Attention-based model generates high-quality images of Indian Sign Language letters, numbers, and words, outperforming the traditional ProGAN in Inception Score (IS) and Fréchet Inception Distance (FID), with improvements of 3.2 and 30.12, respectively. Additionally, we are publishing a large dataset incorporating high-quality images of Indian Sign Language alphabets, numbers, and 129 words.

Abstract (translated)

手语包含手势、面部表情和身体动作,是与听力障碍人士沟通的重要媒介。经过良好训练的手语社区能够轻松交流,但对于不懂手语的人来说则面临巨大挑战。识别和生成是听觉健全者和听力障碍者之间基本的交流方式。尽管在识别方面取得了进展,但手语生成仍需进一步探索。 渐进式生长生成对抗网络(ProGAN)擅长生成高质量图像,而自注意力生成对抗网络(SAGAN)则擅长在中等分辨率下生成细节丰富的图像。对于手语图像生成而言,在平衡分辨率和细节度上至关重要。我们正在开发一种结合这两种模型的生成对抗网络(GAN)变体,以生成特征丰富、高分辨率且具有类别条件的手语图像。 我们的基于注意力机制的改良模型能够生成高质量的印度手语字母、数字及词汇图像,在Inception分数(IS)和Fréchet Inception距离(FID)指标上超越了传统的ProGAN,分别提高了3.2和30.12。此外,我们还计划发布一个包含高分辨率印度手语字母、数字以及129个词语的大规模数据集。

URL

https://arxiv.org/abs/2508.09522

PDF

https://arxiv.org/pdf/2508.09522.pdf


Tags
3D Action Action_Localization Action_Recognition Activity Adversarial Agent Attention Autonomous Bert Boundary_Detection Caption Chat Classification CNN Compressive_Sensing Contour Contrastive_Learning Deep_Learning Denoising Detection Dialog Diffusion Drone Dynamic_Memory_Network Edge_Detection Embedding Embodied Emotion Enhancement Face Face_Detection Face_Recognition Facial_Landmark Few-Shot Gait_Recognition GAN Gaze_Estimation Gesture Gradient_Descent Handwriting Human_Parsing Image_Caption Image_Classification Image_Compression Image_Enhancement Image_Generation Image_Matting Image_Retrieval Inference Inpainting Intelligent_Chip Knowledge Knowledge_Graph Language_Model LLM Matching Medical Memory_Networks Multi_Modal Multi_Task NAS NMT Object_Detection Object_Tracking OCR Ontology Optical_Character Optical_Flow Optimization Person_Re-identification Point_Cloud Portrait_Generation Pose Pose_Estimation Prediction QA Quantitative Quantitative_Finance Quantization Re-identification Recognition Recommendation Reconstruction Regularization Reinforcement_Learning Relation Relation_Extraction Represenation Represenation_Learning Restoration Review RNN Robot Salient Scene_Classification Scene_Generation Scene_Parsing Scene_Text Segmentation Self-Supervised Semantic_Instance_Segmentation Semantic_Segmentation Semi_Global Semi_Supervised Sence_graph Sentiment Sentiment_Classification Sketch SLAM Sparse Speech Speech_Recognition Style_Transfer Summarization Super_Resolution Surveillance Survey Text_Classification Text_Generation Time_Series Tracking Transfer_Learning Transformer Unsupervised Video_Caption Video_Classification Video_Indexing Video_Prediction Video_Retrieval Visual_Relation VQA Weakly_Supervised Zero-Shot