4 papers
UniAV: Unified Audio-Visual Perception for Multi-Task Video Event Localization
Tiantian Geng, Teng Wang, Jinming Duan +4
Video event localization tasks include temporal action localization (TAL), sound event detection (SED) and audio-visual event localization (AVEL). Existing methods tend to over-spe…
HCNQA: Enhancing 3D VQA with Hierarchical Concentration Narrowing Supervision
Shengli Zhou, Jianuo Zhu, Qilin Huang +3
3D Visual Question-Answering (3D VQA) is pivotal for models to perceive the physical world and perform spatial reasoning. Answer-centric supervision is a commonly used training met…
RoboReflect: A Robotic Reflective Reasoning Framework for Grasping Ambiguous-Condition Objects
Zhen Luo, Yixuan Yang, Yanfu Zhang +1
As robotic technology rapidly develops, robots are being employed in an increasing number of fields. However, due to the complexity of deployment environments or the prevalence of…
A Self-guided Multimodal Approach to Enhancing Graph Representation Learning for Alzheimer's Diseases
Zhepeng Wang, Runxue Bao, Yawen Wu +6
Graph neural networks (GNNs) are powerful machine learning models designed to handle irregularly structured data. However, their generic design often proves inadequate for analyzin…