InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring

· Source: Artificial Intelligence · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Robotics & Autonomous Systems · Depth: Expert, quick

Summary

InCarEmo is a new multimodal dataset designed for in-cabin emotion recognition and driver state monitoring, addressing limitations in existing public datasets that primarily rely on visual modalities and lack conversational information. This dataset integrates RGB and infrared video, in-cabin audio, and dialogue text, collected from scripted scenarios simulating realistic driver behaviors under diverse lighting and driving contexts. InCarEmo supports multimodal emotion recognition, fatigue detection, and distraction monitoring. It includes an auxiliary English benchmark for cross-lingual evaluation and provides a unified benchmark with extensive baseline results across unimodal and multimodal methods, analyzing performance under modality-missing and noise conditions. Experimental results highlight the benefits of multimodal fusion while revealing challenges in real-world noise and low-light environments.

Key takeaway

For AI scientists and engineers developing next-generation intelligent in-cabin systems, recognize that relying solely on visual data is insufficient for robust driver state monitoring. You should integrate multimodal inputs, including conversational audio and text, alongside visual cues to capture the full spectrum of driver emotion and state. Utilize datasets like InCarEmo to train and benchmark models, focusing on improving performance under challenging real-world noise and low-light conditions to enhance safety and human-vehicle interaction.

Key insights

InCarEmo is a multimodal dataset integrating visual, audio, and text data to advance in-cabin emotion and driver state understanding.

Principles

Method

InCarEmo was constructed using scripted in-cabin scenarios, collecting RGB/infrared video, audio, and dialogue text, then benchmarked with unimodal and multimodal methods under various conditions.

In practice

Topics

Best for: Research Scientist, AI Scientist, Computer Vision Engineer, NLP Engineer

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Artificial Intelligence.