2 papers
cs.MM2026
Dual-Stream Decoupled Learning for Temporal Consistency and Speaker Interaction in AVSD
Junhao Xiao, Shun Feng, Zhiyu Wu +6
Audio-Visual Speaker Detection (AVSD) hinges on modeling both individual temporal continuity and inter-personal social context. Existing coupled architectures struggle to reconcile…
cs.CV2026
Not Just What's There: Enabling CLIP to Comprehend Negated Visual Descriptions Without Fine-tuning
Junhao Xiao, Zhiyu Wu, Hao Lin +5
Vision-Language Models (VLMs) like CLIP struggle to understand negation, often embedding affirmatives and negatives similarly (e.g., matching "no dog" with dog images). Existing me…