12 papers
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs
Yuhan Zhu, Changlian Ma, Xiangyu Zeng +12
Video multimodal large language models (MLLMs) can describe what happens in a video, but rarely identify when the supporting evidence occurs. We study generalist video temporal gro…
EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning
Zhitong Wang, Songze Li, Hao Peng +4
Reinforcement learning (RL) has emerged as a powerful paradigm for training Large Language Models (LLMs) as agents. However, conventional RL methods for long-horizon agentic tasks…
Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction
Tianxiang Jiang, Linquan Wu, Sheng Xia +5
Video event prediction (VEP) requires models to infer unobserved future states from partial video evidence. Existing video MLLMs usually verbalize intermediate future reasoning in…
Delta-Adapter: Scalable Exemplar-Based Image Editing with Single-Pair Supervision
Jiacheng Chen, Songze Li, Han Fu +5
Exemplar-based image editing applies a transformation defined by a source-target image pair to a new query image. Existing methods rely on a pair-of-pairs supervision paradigm, req…
Cross-Modal Backdoors in Multimodal Large Language Models
Runhe Wang, Li Bai, Haibo Hu +1
Developers increasingly construct multimodal large language models (MLLMs) by assembling pretrained components,introducing supply-chain attack surfaces.Existing security research p…
Learning Goal-Oriented Vision-and-Language Navigation with Self-Improving Demonstrations at Scale
Songze Li, Zun Wang, Gengze Zhou +8
Goal-oriented vision-language navigation requires robust exploration capabilities for agents to navigate to specified goals in unknown environments without step-by-step instruction…