works on

From the 1 of 17 linked papers with an AI index.

collaborators

17 papers

cs.CV2026

The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering

Yuqian Fu, Tianwen Qian, Yanjun Li +30

EgoCross is a cross-domain egocentric video question answering benchmark designed to evaluate whether multimodal large language models can generalize beyond common daily-life scena…

cs.CV2026

ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding

Mingkang Dong, Muxin Pu, Jie Li +8

ObjectStream introduces a training‑free method that extracts latent objects from frozen Video‑LLM representations and uses them as persistent memory anchors to improve streaming vi…

cs.LG2026

3SPO: State-Score-Supervised Policy Optimization for LLM Agents

Yu Han, Kailing Li, Yang Jiao +4

Training large language models (LLMs) as autonomous agents via reinforcement learning (RL) has enabled frontier models to achieve superhuman performance in long-horizon tasks. Howe…

cs.CV2026

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning

Kailing Li, Qi'ao Xu, Tianwen Qian +3

Embodied Visual Reasoning (EVR) seeks to follow complex, free-form instructions based on egocentric video, enabling semantic understanding and spatiotemporal reasoning in dynamic e…

cs.RO2026

Afford-VLA: Action-Aligned Visual Planning via Internalized Affordance

Runze Wang, Yuqian Fu, Yu Li +7

Vision-language-action (VLA) models have shown strong potential for generalist robot manipulation, yet they remain limited by insufficient spatial reasoning, particularly in determ…

cs.CV2026

SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning

Yian Li, Yang Jiao, Bin Zhu +4

Spatial intelligence, which refers to the ability to reason about geometric and physical structure from visual observations, remains a core challenge for multimodal large language…