#multimodal dataset

topicmultimodal dataset

10 papers · 1 filter

cs.CV2026

ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine

Yukang Cao, Haozhe Xie, Beichen Wen +13

The paper presents ACE, a data collection system that records synchronized multimodal streams—including egocentric and multi-view video, full-body and hand motion, object geometry,…

cs.CV2026

Theia: Large-Scale Multimodal Captioning and Automated Validation of the Incidents1M Dataset for Data-Free Distillation

Simone Giano, Lorenzo Severini, Alessandro Galdelli +1

The paper presents a method to create and automatically validate a large multimodal disaster‑response dataset by generating captions for vision‑only images using LLMs and evaluatin…

cs.SD2026

SKY-Piano: A Multimodal Piano Performance Dataset

Joonhyung Bae, Dawon Park, Taegyun Kwon +10

The paper introduces SKY-Piano, a multimodal dataset of piano performances that includes audio, MIDI, multi-view video, hand and body motion capture, and MusicXML scores from profe…

eess.AS2026

CARE: A Multimodal Corpus for Studying Speech and Non-Verbal Communication Across Multiple Medical Conditions

David Gimeno-Gómez, Catarina Botelho, Carlos-D. Martínez-Hinarejos +2

The paper introduces CARE v1.0, a curated multimodal English dataset of about 144 hours of short video interviews from 612 participants covering 12 medical conditions and a control…

cs.AI2026

InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring

Hao Yang, Yanyan Zhao, Kewei Zhao +11

The paper presents InCarEmo, a multimodal dataset that combines RGB and infrared video, audio, and dialogue text for in-cabin emotion recognition, fatigue detection, and distractio…

cs.HC2026

VIP-MINGLE: A Corpus for Videoconference and In-Person Multimodal Interaction in Group Language Engagement

Andrew Chang, Abhinay K Bodi, Wenxin Deng +6

The paper presents VIP-MINGLE, a multimodal dataset of group conversations recorded both in videoconference and in-person settings, including audio, video, facial expressions, tran…