12 papers
Geometry-Grounded Unified 3D Perception for Autonomous Driving
Longfei Xu, Xiaohui Wang, Zehao Huang +4
Camera-based autonomous driving perception requires a shared representation that preserves metric 3D structure across synchronized multi-camera streams. However, existing image-bas…
MoCA: Implicit Social Context Analysis
Wenhao Xu, Kaiwen Zhang, Hao Li +10
Human social communication, such as affection and intent, is often conveyed in highly implicit ways, where underlying meanings are expressed through indirect, socially and cultural…
From Profiling to Synthesis: Benchmarking Implicit Behavioral Alignment in Personalized LLM Agents
Jiajia Song, Bobo Li, Haiwen Yi +6
Large Language Models have enabled increasingly capable autonomous agents, yet personalization remains critical for making such agents practically useful. Recent benchmarks have be…
Audio-Visual Intelligence in Large Foundation Models
You Qin, Kai Liu, Shengqiong Wu +12
Audio-Visual Intelligence (AVI) has emerged as a central frontier in artificial intelligence, bridging auditory and visual modalities to enable machines that can perceive, generate…
LASQ: A Low-resource Aspect-based Sentiment Quadruple Extraction Dataset
Aizihaierjiang Yusufu, Jiang Liu, Kamran Aziz +5
In recent years, aspect-based sentiment analysis (ABSA) has made rapid progress and shown strong practical value. However, existing research and benchmarks are largely concentrated…
Unveiling the Cognitive Compass: Theory-of-Mind-Guided Multimodal Emotion Reasoning
Meng Luo, Bobo Li, Shanqing Xu +8
Despite rapid progress in multimodal large language models (MLLMs), their capability for deep emotional understanding remains limited. We argue that genuine affective intelligence…