From the 1 of 5 linked papers with an AI index.
5 papers
InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring
Hao Yang, Yanyan Zhao, Kewei Zhao +11
The paper presents InCarEmo, a multimodal dataset that combines RGB and infrared video, audio, and dialogue text for in-cabin emotion recognition, fatigue detection, and distractio…
DeepSight: Bridging Depth Maps and Language with a Depth-Driven Multimodal Model
Hao Yang, Hongbo Zhang, Yanyan Zhao +1
Multimodal large language models (MLLMs) have achieved impressive performance across various tasks such as image captioning and visual question answer(VQA); however, they often str…
Text-Driven Emotionally Continuous Talking Face Generation
Hao Yang, Yanyan Zhao, Tian Zheng +7
Talking Face Generation (TFG) strives to create realistic and emotionally expressive digital faces. While previous TFG works have mastered the creation of naturalistic facial movem…
CARE-Bench: A Benchmark of Diverse Client Simulations Guided by Expert Principles for Evaluating LLMs in Psychological Counseling
Bichen Wang, Yixin Sun, Junzhe Wang +6
The mismatch between the growing demand for psychological counseling and the limited availability of services has motivated research into the application of Large Language Models (…
Data Uncertainty-Aware Learning for Multimodal Aspect-based Sentiment Analysis
Hao Yang, Zhenyu Zhang, Yanyan Zhao +1
As a fine-grained task, multimodal aspect-based sentiment analysis (MABSA) mainly focuses on identifying aspect-level sentiment information in the text-image pair. However, we obse…