works on

From the 2 of 12 linked papers with an AI index.

collaborators

12 papers

cs.CV2026

GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience

Sitong Gong, Caixin Kang, Tianyu Yan +7

A wearable assistant should both answer questions about its visual history and recognize when that history is useful to the present situation. Existing video-memory systems primari…

cs.IR2026

SQuTR: A Robustness Benchmark for Spoken Query to Text Retrieval under Acoustic Noise

Yuejie Li, Ke Yang, Yueying Hua +4

The paper introduces SQuTR, a benchmark dataset and evaluation protocol for testing how well spoken query‑to‑text retrieval systems perform under various levels of real‑world acous…

cs.CV2026

Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos

Gong Sitong, Tianyu Yan, Caixin Kang +6

The paper introduces Vinci2, a proactive on‑device assistant for continuous egocentric video that decides when to intervene by using memory‑augmented reasoning, and presents EgoSer…

cs.CV2026

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering

Mingfang Zhang, Jingjing Pan, Ashutosh Kumar +7

Cause-and-effect reasoning in video is a significant challenge for Vision-Language Models (VLMs), as it requires going beyond surface-level perception to a deeper understanding of…

cs.AI2026

Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?

Caixin Kang, Tianyu Yan, Sitong Gong +8

Multimodal Large Language Models (MLLMs) are increasingly deployed in human-facing roles where personality perception is critical, yet existing benchmarks evaluate this capability…

cs.CV2026

SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting

Ruicong Liu, Yifei Huang, Liangyang Ouyang +2

Real-time 3D hand forecasting is a critical component for fluid human-computer interaction in applications like AR and assistive robotics. However, existing methods are ill-suited…