4
- https://doi.org/10.1109/icarc64760.2025.10962991
SSL400 - A Comprehensive Word Level Dataset for Sinhala Sign Language Recognition
- Feb 19, 2025
- Yohan Abhishek +1 more
In the context of increasing digital accessibility and the need for inclusive communication technologies, the demand for automated sign language recognition has grown significantly. However, Sinhala Sign Language (SSL), used by the Sri Lankan deaf community, remains underexplored due to the lack of labelled datasets and the challenges of developing recognition models for low-resource languages. This study addresses this gap by introducing an SSL dataset, specifically curated to support real-time recognition of 384 commonly used words in daily and work-related contexts. The methodology for dataset creation involved expert consultations, a structured data collection process, and multi-step labelling. These steps ensured the accurate labeling of video data, which can be used for machine learning tasks like gesture recognition, sentence formation, and speech translation. The proposed dataset lays the foundation for SSL recognition models that are fast, lightweight, and optimized for low-resource environments. Additionally, we explore the performance of initial recognition models using this dataset, demonstrating its applicability in real-world scenarios. The dataset and research outcomes provide valuable resources for the development of inclusive AI-driven communication tools for the Sinhala-speaking deaf community. This contribution addresses the pressing need for better technological support for underrepresented languages in the field of sign language recognition and processing.