collaborators

12 papers

cs.CV2026

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs

Yuhan Zhu, Changlian Ma, Xiangyu Zeng +12

Video multimodal large language models (MLLMs) can describe what happens in a video, but rarely identify when the supporting evidence occurs. We study generalist video temporal gro…

cs.LG2026

EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning

Zhitong Wang, Songze Li, Hao Peng +4

Reinforcement learning (RL) has emerged as a powerful paradigm for training Large Language Models (LLMs) as agents. However, conventional RL methods for long-horizon agentic tasks…

cs.CV2026

Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction

Tianxiang Jiang, Linquan Wu, Sheng Xia +5

Video event prediction (VEP) requires models to infer unobserved future states from partial video evidence. Existing video MLLMs usually verbalize intermediate future reasoning in…

cs.CV2026

Delta-Adapter: Scalable Exemplar-Based Image Editing with Single-Pair Supervision

Jiacheng Chen, Songze Li, Han Fu +5

Exemplar-based image editing applies a transformation defined by a source-target image pair to a new query image. Existing methods rely on a pair-of-pairs supervision paradigm, req…

cs.CR2026

Cross-Modal Backdoors in Multimodal Large Language Models

Runhe Wang, Li Bai, Haibo Hu +1

Developers increasingly construct multimodal large language models (MLLMs) by assembling pretrained components,introducing supply-chain attack surfaces.Existing security research p…

cs.CV2026

Learning Goal-Oriented Vision-and-Language Navigation with Self-Improving Demonstrations at Scale

Songze Li, Zun Wang, Gengze Zhou +8

Goal-oriented vision-language navigation requires robust exploration capabilities for agents to navigate to specified goals in unknown environments without step-by-step instruction…