collaborators

6 papers

cs.CV2025

EgoBlind: Towards Egocentric Visual Assistance for the Blind

Junbin Xiao, Nanxin Huang, Hao Qiu +5

We present EgoBlind, the first egocentric VideoQA dataset collected from blind individuals to evaluate the assistive capabilities of contemporary multimodal large language models (…

cs.CV2025

Scene-Text Grounding for Text-Based Video Question Answering

Sheng Zhou, Junbin Xiao, Xun Yang +5

Existing efforts in text-based video question answering (TextVideoQA) are criticized for their opaque decisionmaking and heavy reliance on scene-text recognition. In this paper, we…

cs.CV2025

Towards Efficient Partially Relevant Video Retrieval with Active Moment Discovering

Peipei Song, Long Zhang, Long Lan +4

Partially relevant video retrieval (PRVR) is a practical yet challenging task in text-to-video retrieval, where videos are untrimmed and contain much background content. The pursui…

cs.CV2025

Unified Static and Dynamic Network: Efficient Temporal Filtering for Video Grounding

Jingjing Hu, Dan Guo, Kun Li +4

Inspired by the activity-silent and persistent activity mechanisms in human visual perception biology, we design a Unified Static and Dynamic Network (UniSDNet), to learn the seman…

cs.CV2025

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering

Sheng Zhou, Junbin Xiao, Qingyun Li +6

We introduce EgoTextVQA, a novel and rigorously constructed benchmark for egocentric QA assistance involving scene text. EgoTextVQA contains 1.5K ego-view videos and 7K scene-text…

cs.CV2025

A Survey on fMRI-based Brain Decoding for Reconstructing Multimodal Stimuli

Pengyu Liu, Guohua Dong, Dan Guo +5

In daily life, we encounter diverse external stimuli, such as images, sounds, and videos. As research in multimodal stimuli and neuroscience advances, fMRI-based brain decoding has…