2 citations · 3 across the 6 of their papers we have counts for
9 papers
SSAM: Self-Supervised Association Modeling for Test-Time Adaption
Yaxiong Wang, Zhenqiang Zhang, Lechao Cheng +3
Test-time adaption (TTA) has witnessed important progress in recent years, the prevailing methods typically first encode the image and the text and design strategies to model the a…
Towards Efficient Partially Relevant Video Retrieval with Active Moment Discovering
Peipei Song, Long Zhang, Long Lan +4
Partially relevant video retrieval (PRVR) is a practical yet challenging task in text-to-video retrieval, where videos are untrimmed and contain much background content. The pursui…
EmoSEM: Segment and Explain Emotion Stimuli in Visual Art
Jing Zhang, Dan Guo, Zhangbin Li +1
This paper focuses on a key challenge in visual emotion understanding: given an art image, the model pinpoints pixel regions that trigger a specific human emotion, and generates li…
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
Sheng Zhou, Junbin Xiao, Qingyun Li +6
We introduce EgoTextVQA, a novel and rigorously constructed benchmark for egocentric QA assistance involving scene text. EgoTextVQA contains 1.5K ego-view videos and 7K scene-text…
Repetitive Action Counting with Hybrid Temporal Relation Modeling
Kun Li, Xinge Peng, Dan Guo +2
Repetitive Action Counting (RAC) aims to count the number of repetitive actions occurring in videos. In the real world, repetitive actions have great diversity and bring numerous c…
MOL-Mamba: Enhancing Molecular Representation with Structural & Electronic Insights
Jingjing Hu, Dan Guo, Zhan Si +5
Molecular representation learning plays a crucial role in various downstream tasks, such as molecular property prediction and drug design. To accurately represent molecules, Graph…