4 papers
YOLOA: Real-Time Affordance Detection via LLM Adapter
Yuqi Ji, Junjie Ke, Lihuo He +5
Affordance detection aims to jointly address the fundamental "what-where-how" challenge in embodied AI by understanding "what" an object is, "where" the object is located, and "how…
Boosting Temporal Sentence Grounding via Causal Inference
Kefan Tang, Lihuo He, Jisheng Dang +1
Temporal Sentence Grounding (TSG) aims to identify relevant moments in an untrimmed video that semantically correspond to a given textual query. Despite existing studies having mad…
EyeSim-VQA: A Free-Energy-Guided Eye Simulation Framework for Video Quality Assessment
Zhaoyang Wang, Wen Lu, Jie Li +3
Free-energy-guided self-repair mechanisms have shown promising results in image quality assessment (IQA), but remain under-explored in video quality assessment (VQA), where tempora…
A Multi-annotated and Multi-modal Dataset for Wide-angle Video Quality Assessment
Bo Hu, Wei Wang, Chunyi Li +3
Wide-angle video is favored for its wide viewing angle and ability to capture a large area of scenery, making it an ideal choice for sports and adventure recording. However, wide-a…