3 papers
cs.CV2026
RA-SSU: Towards Fine-Grained Audio-Visual Learning with Region-Aware Sound Source Understanding
Muyi Sun, Yixuan Wang, Hong Wang +5
Audio-Visual Learning (AVL) is one fundamental task of multi-modality learning and embodied intelligence, displaying the vital role in scene understanding and interaction. However,…
cs.CV2024
Benchmarking Badminton Action Recognition with a New Fine-Grained Dataset
Qi Li, Tzu-Chen Chiu, Hsiang-Wei Huang +2
In the dynamic and evolving field of computer vision, action recognition has become a key focus, especially with the advent of sophisticated methodologies like Convolutional Neural…
cs.CV2024
Ego3DT: Tracking Every 3D Object in Ego-centric Videos
Shengyu Hao, Wenhao Chai, Zhonghan Zhao +8
The growing interest in embodied intelligence has brought ego-centric perspectives to contemporary research. One significant challenge within this realm is the accurate localizatio…