The In-Car Sign Language Corpus (ICSL): A Multi-Modal Resource for Constrained-Space Sign Language Recognition

· Source: Takara TLDR - Daily AI Papers · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Data Science & Analytics · Depth: Expert, medium

Summary

The In-Car Sign Language (ICSL) dataset is introduced as a multi-modal resource to address the challenges of sign language recognition (SLR) within confined vehicle interiors, specifically for Brazilian Sign Language (Libras). This corpus aims to enhance public transport accessibility for the Deaf and Hard-of-Hearing community. The dataset comprises over 1.5 million frames, featuring two main components: high-precision laboratory motion capture (MoCap) data for an idealized linguistic baseline, and real-world multi-modal in-car recordings captured using 2D cameras and 3D Time-of-Flight sensors. These synchronized streams, collected from Libras users in various in-car scenarios, are provided with gloss annotations for both lexical and non-lexical sign language elements. ICSL supports comparative analyses between synthesized signing avatars and real interpreter videos, facilitating research into robust "in-the-wild" SLR models and domain adaptation for constrained, occluded, and non-frontal environments.

Key takeaway

For Machine Learning Engineers developing sign language recognition systems for real-world applications, you should consider the unique challenges of constrained environments like vehicle interiors. The ICSL dataset offers a critical resource for training robust models, providing both idealized MoCap and multi-modal in-car data. Utilize this corpus to develop and evaluate deep neural networks, specifically addressing occlusion and non-frontal signing, to enhance accessibility for Deaf and Hard-of-Hearing communities in shared mobility.

Key insights

The ICSL dataset enables sign language recognition research in challenging, confined vehicle environments.

Principles

Method

The ICSL dataset collection protocol involves high-precision lab MoCap and real-world in-car recordings using 2D cameras and 3D Time-of-Flight sensors, synchronized and gloss-annotated for deep neural network training.

In practice

Topics

Code references

Best for: Research Scientist, NLP Engineer, AI Scientist, Machine Learning Engineer, Computer Vision Engineer

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Takara TLDR - Daily AI Papers.