16 papers
Bridging the Agent-World Gap: Text World Models for LLM-based Agents
Yixia Li, Hongru Wang, Peng Lai +13
Large language model (LLM)-based agents are increasingly used in interactive textual environments, from web navigation and code editing to tool use and long-horizon dialogue. Yet m…
Self-Prophetic Decoding to Unlock Visual Search in LVLMs
Zhendong He, Qiyuan Dai, Guanbin Li +2
Large Vision-Language Models (LVLMs) are rapidly evolving toward true multimodal reasoning, with visual search representing a concrete instantiation of the thinking-with-images par…
EVA: Editing for Versatile Alignment against Jailbreaks
Yi Wang, Hongye Qiu, Yue Xu +4
Large Language Models (LLMs) and Vision Language Models (VLMs) have demonstrated impressive capabilities but remain vulnerable to jailbreaking attacks, where adversaries exploit te…
Vision Transformers Need More Than Registers
Cheng Shi, Yizhou Yu, Sibei Yang
Vision Transformers (ViTs), when pre-trained on large-scale data, provide general-purpose representations for diverse downstream tasks. However, artifacts in ViTs are widely observ…
Chart Deep Research in LVLMs via Parallel Relative Policy Optimization
Jiajin Tang, Gaoyang, Wenjie Wang +2
With the rapid advancement of data science, charts have evolved from simple numerical presentation tools to essential instruments for insight discovery and decision-making support.…
WeaveTime: Stream from Earlier Frames into Emergent Memory in VideoLLMs
Yulin Zhang, Cheng Shi, Sibei Yang
Recent advances in Multimodal Large Language Models have greatly improved visual understanding and reasoning, yet their quadratic attention and offline training protocols make them…