9 papers
VSAS-Bench: Real-Time Evaluation of Visual Streaming Assistant Models
Pavan Kumar Anasosalu Vasu, Cem Koc, Fartash Faghri +6
Streaming vision-language models (VLMs) continuously generate responses given an instruction prompt and an online stream of input frames. This is a core mechanism for real-time vis…
VERDI: VLM-Embedded Reasoning for Autonomous Driving
Bowen Feng, Zhiting Mei, Julian Ost +5
While autonomous driving (AD) stacks struggle with decision making under partial observability and real-world complexity, human drivers are capable of applying commonsense reasonin…
When Large Language Models are More PersuasiveThan Incentivized Humans, and Why
Philipp Schoenegger, Francesco Salvi, Jiacheng Liu +39
Large Language Models (LLMs) have been shown to be highly persuasive, but when and why they outperform humans is still an open question. We compare the persuasiveness of two LLMs (…
Training LLMs Beyond Next Token Prediction -- Filling the Mutual Information Gap
Chun-Hao Yang, Bo-Han Feng, Tzu-Yuan Lai +3
Optimizing training performance in large language models (LLMs) remains an essential challenge, particularly in improving model performance while maintaining computational costs. T…
Alignment Without Understanding: A Message- and Conversation-Centered Approach to Understanding AI Sycophancy
Lihua Du, Xing Lyu, Lezi Xie +1
AI sycophancy is increasingly recognized as a harmful alignment, but research remains fragmented and underdeveloped at the conceptual level. This article redefines AI sycophancy as…
StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming Assistant
Haibo Wang, Bo Feng, Zhengfeng Lai +6
We present StreamBridge, a simple yet effective framework that seamlessly transforms offline Video-LLMs into streaming-capable models. It addresses two fundamental challenges in ad…