8 citations · 11 across the 13 of their papers we have counts for
10 papers · 1 filter
Disentangling Foreground and Background for vision-Language Navigation via Online Augmentation
Yunbo Xu, Xuesong Zhang, Jia Li +2
Following language instructions, vision-language navigation (VLN) agents are tasked with navigating unseen environments. While augmenting multifaceted visual representations has pr…
Generalizable Engagement Estimation in Conversation via Domain Prompting and Parallel Attention
Yangche Yu, Yin Chen, Jia Li +6
Accurate engagement estimation is essential for adaptive human-computer interaction systems, yet robust deployment is hindered by poor generalizability across diverse domains and c…
Rebalancing Contrastive Alignment with Bottlenecked Semantic Increments in Text-Video Retrieval
Jian Xiao, Zijie Song, Jialong Hu +4
Recent progress in text-video retrieval has been largely driven by contrastive learning. However, existing methods often overlook the effect of the modality gap, which causes ancho…
VAEmo: Efficient Representation Learning for Visual-Audio Emotion with Knowledge Injection
Hao Cheng, Zhiwei Zhao, Yichao He +4
Audiovisual emotion recognition (AVER) aims to infer human emotions from nonverbal visual-audio (VA) cues, offering modality-complementary and language-agnostic advantages. However…
Video Flow as Time Series: Discovering Temporal Consistency and Variability for VideoQA
Zijie Song, Zhenzhen Hu, Yixiao Ma +2
Video Question Answering (VideoQA) is a complex video-language task that demands a sophisticated understanding of both visual content and temporal dynamics. Traditional Transformer…
PhysioSync: Temporal and Cross-Modal Contrastive Learning Inspired by Physiological Synchronization for EEG-Based Emotion Recognition
Kai Cui, Jia Li, Yu Liu +3
Electroencephalography (EEG) signals provide a promising and involuntary reflection of brain activity related to emotional states, offering significant advantages over behavioral c…