works on

From the 1 of 46 linked papers with an AI index.

activity
20242026
collaborators

46 papers

cs.CV2026

Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents

Wenqi Liu, Shijie Ma, Yunxiao Wang +21

Open-world video understanding often requires a model to locate sparse visual evidence and acquire external knowledge that is absent from the video and its parametric memory. While…

cs.CL2026

From Reasoning Depth to Reasoning Breadth: Evaluating Multi-Point Associative Reasoning in Large Language Models

Si'an Xie, Jiaxun Liu, Biao Yang +4

Large language models (LLMs) have made substantial progress on reasoning tasks that require increasingly long and complex inferential chains. This progress primarily reflects reaso…

cs.CV2026

LaME: Learning to Think in Latent Space for Multimodal Embedding via Information Bottleneck

Peixi Wu, Biao Yang, Feipeng Ma +7

The paper introduces LaME, a multimodal embedding model that performs reasoning in a compact latent space using learnable tokens and an information‑bottleneck objective, eliminatin…

cs.CL2026

PraMem: Practice-derived Experiential Memory for Long-horizon Behavior Prediction

Zhuoqun Li, Boxi Cao, Jiawei Chen +11

Long-horizon behavior prediction aims to infer a user's next action based on a lengthy historical sequence, playing a crucial role in artificial intelligence field. The rise of lar…

cs.CV2026

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences

Yankai Yang, Yancheng Long, Bin Wen +4

Video multimodal large language models have made strong progress on open-ended video understanding, but they still lack precise local spatiotemporal perception. When two videos sha…

cs.CV2026

SpatialFlow-GRPO: Where Spatial Credit Drives Image Editing

Yankai Yang, Yancheng Long, Wei Chen +7

Recent online reinforcement learning has substantially improved image editing quality. However, existing Flow-GRPO-style methods usually rely on a single whole-image reward, which…