From the 1 of 7 linked papers with an AI index.
7 papers
ReToken: One Token to Improve Vision-Language Models for Visual Retrieval
Yao Xiao, Reuben Tan, Zhen Zhu +3
ReToken introduces a single learnable embedding that acts as a retrieval token to select a sparse set of relevant visual tokens from a cached representation, improving vision-langu…
All-in-One Video Restoration under Smoothly Evolving Unknown Weather Degradations
Wenrui Li, Hongtao Chen, Yao Xiao +4
All-in-one image restoration aims to recover clean images from diverse unknown degradations using a single model. But extending this task to videos faces unique challenges. Existin…
Sparse Reasoning is Enough: Biological-Inspired Framework for Video Anomaly Detection with Large Pre-trained Models
He Huang, Zixuan Hu, Dongxiao Li +2
Video anomaly detection (VAD) plays a vital role in real-world applications such as security surveillance, autonomous driving, and industrial monitoring. Recent advances in large p…
TextRegion: Text-Aligned Region Tokens from Frozen Image-Text Models
Yao Xiao, Qiqian Fu, Heyi Tao +3
Image-text models excel at image-level tasks but struggle with detailed visual understanding. While these models provide strong visual-language alignment, segmentation models like…
Alzheimer's Dementia Detection Using Perplexity from Paired Large Language Models
Yao Xiao, Heidi Christensen, Stefan Goetze
Alzheimer's dementia (AD) is a neurodegenerative disorder with cognitive decline that commonly impacts language ability. This work extends the paired perplexity approach to detecti…
Early Dementia Detection Using Multiple Spontaneous Speech Prompts: The PROCESS Challenge
Fuxiang Tao, Bahman Mirheidari, Madhurananda Pahar +12
Dementia is associated with various cognitive impairments and typically manifests only after significant progression, making intervention at this stage often ineffective. To addres…