Paper Reading AI Learner

Indian Sign Language Detection for Real-Time Translation using Machine Learning

2025-07-31 08:51:49
Rajat Singhal, Jatin Gupta, Akhil Sharma, Anushka Gupta, Navya Sharma

Abstract

Gestural language is used by deaf & mute communities to communicate through hand gestures & body movements that rely on visual-spatial patterns known as sign languages. Sign languages, which rely on visual-spatial patterns of hand gestures & body movements, are the primary mode of communication for deaf & mute communities worldwide. Effective communication is fundamental to human interaction, yet individuals in these communities often face significant barriers due to a scarcity of skilled interpreters & accessible translation technologies. This research specifically addresses these challenges within the Indian context by focusing on Indian Sign Language (ISL). By leveraging machine learning, this study aims to bridge the critical communication gap for the deaf & hard-of-hearing population in India, where technological solutions for ISL are less developed compared to other global sign languages. We propose a robust, real-time ISL detection & translation system built upon a Convolutional Neural Network (CNN). Our model is trained on a comprehensive ISL dataset & demonstrates exceptional performance, achieving a classification accuracy of 99.95%. This high precision underscores the model's capability to discern the nuanced visual features of different signs. The system's effectiveness is rigorously evaluated using key performance metrics, including accuracy, F1 score, precision & recall, ensuring its reliability for real-world applications. For real-time implementation, the framework integrates MediaPipe for precise hand tracking & motion detection, enabling seamless translation of dynamic gestures. This paper provides a detailed account of the model's architecture, the data preprocessing pipeline & the classification methodology. The research elaborates the model architecture, preprocessing & classification methodologies for enhancing communication in deaf & mute communities.

Abstract (translated)

手势语言是由聋哑社区通过手部动作和身体姿态的视觉空间模式来交流的一种方式,这种模式被称为手语。手语依靠手部动作和身体姿态的视觉空间模式,是全球聋哑人群的主要沟通方式。有效的沟通对人类互动至关重要,然而由于熟练翻译人员和可访问翻译技术稀缺,这些群体常面临重大障碍。 本研究针对印度语境下的这些问题,特别关注印度手语(ISL),通过机器学习来填补听障人士的交流缺口。在印度,相对于全球其他手语来说,ISL的技术解决方案发展相对滞后。我们提出了一种基于卷积神经网络(CNN)的强大实时ISL检测与翻译系统。我们的模型经过全面的ISL数据集训练,并展示了卓越性能,达到了99.95%的分类准确率。这一高精度表明了该模型能够识别不同手势之间的细微视觉特征差异。 为了评估系统的有效性,我们使用了一系列关键性能指标进行严格测试,包括准确性、F1分数、精确度和召回率,确保其在实际应用中的可靠性。对于实时实施,框架整合了MediaPipe以实现精准的手部追踪及运动检测,使得动态手势的无缝翻译得以实现。 本文详细描述了模型架构、数据预处理管道以及分类方法,并深入探讨了增强聋哑人群体沟通的方法和策略。

URL

https://arxiv.org/abs/2507.20414

PDF

https://arxiv.org/pdf/2507.20414.pdf


Tags
3D Action Action_Localization Action_Recognition Activity Adversarial Agent Attention Autonomous Bert Boundary_Detection Caption Chat Classification CNN Compressive_Sensing Contour Contrastive_Learning Deep_Learning Denoising Detection Dialog Diffusion Drone Dynamic_Memory_Network Edge_Detection Embedding Embodied Emotion Enhancement Face Face_Detection Face_Recognition Facial_Landmark Few-Shot Gait_Recognition GAN Gaze_Estimation Gesture Gradient_Descent Handwriting Human_Parsing Image_Caption Image_Classification Image_Compression Image_Enhancement Image_Generation Image_Matting Image_Retrieval Inference Inpainting Intelligent_Chip Knowledge Knowledge_Graph Language_Model LLM Matching Medical Memory_Networks Multi_Modal Multi_Task NAS NMT Object_Detection Object_Tracking OCR Ontology Optical_Character Optical_Flow Optimization Person_Re-identification Point_Cloud Portrait_Generation Pose Pose_Estimation Prediction QA Quantitative Quantitative_Finance Quantization Re-identification Recognition Recommendation Reconstruction Regularization Reinforcement_Learning Relation Relation_Extraction Represenation Represenation_Learning Restoration Review RNN Robot Salient Scene_Classification Scene_Generation Scene_Parsing Scene_Text Segmentation Self-Supervised Semantic_Instance_Segmentation Semantic_Segmentation Semi_Global Semi_Supervised Sence_graph Sentiment Sentiment_Classification Sketch SLAM Sparse Speech Speech_Recognition Style_Transfer Summarization Super_Resolution Surveillance Survey Text_Classification Text_Generation Time_Series Tracking Transfer_Learning Transformer Unsupervised Video_Caption Video_Classification Video_Indexing Video_Prediction Video_Retrieval Visual_Relation VQA Weakly_Supervised Zero-Shot