From the 1 of 9 linked papers with an AI index.
9 papers
Paths: Prompt-aware Spatio-temporal Transformer with Hierarchical Multi-modal Fusion for RGB-Event Video Person Re-Identification
Yakun Huo, Yingquan Wang, Yangyang Liu +4
RGB-Event Video Person Re-Identification (RE-VReID) aims to retrieve specific person across non-overlapping cameras with complementary RGB videos and event streams. However, existi…
GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience
Sitong Gong, Caixin Kang, Tianyu Yan +7
A wearable assistant should both answer questions about its visual history and recognize when that history is useful to the present situation. Existing video-memory systems primari…
Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos
Gong Sitong, Tianyu Yan, Caixin Kang +6
The paper introduces Vinci2, a proactive on‑device assistant for continuous egocentric video that decides when to intervene by using memory‑augmented reasoning, and presents EgoSer…
3D Segment Anything Model with Visual Mamba for Diagnosing Placenta Accreta Spectrum
Yuliang Zhang, Fang He, Lulu Peng +5
Placenta Accreta Spectrum (PAS) is a rare but highly dangerous obstetric disease. Early and accurate PAS diagnosis is critical for maternal health. Traditional PAS diagnosis relies…
Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?
Caixin Kang, Tianyu Yan, Sitong Gong +8
Multimodal Large Language Models (MLLMs) are increasingly deployed in human-facing roles where personality perception is critical, yet existing benchmarks evaluate this capability…
HFP-SAM: Hierarchical Frequency Prompted SAM for Efficient Marine Animal Segmentation
Pingping Zhang, Tianyu Yan, Yuhao Wang +7
Marine Animal Segmentation (MAS) aims at identifying and segmenting marine animals from complex marine environments. Most of previous deep learning-based MAS methods struggle with…