5 papers
SearchEyes: Towards Frontier Multimodal Deep Search Intelligence via Search World Simulation
Zhengbo Jiao, Yiming Cheng, Yilei Jiang +15
Training multimodal search agents to perform multi-hop reasoning remains challenging due to a fundamental structural disconnect: existing pipelines construct training data, search…
FraudSMSWalker: Benchmarking Agentic Large Language Models for SMS-to-Webpage Fraud Detection
Y. H. Zhou, Z. M. Ma, Y. J. Zhou +12
SMS fraud is increasingly cross-channel: a message directs the user to a webpage, and the final risk depends on how the SMS claim aligns with the page content and requested user ac…
Can MLLMs Reason Beyond Language? VisReason: A Comprehensive Benchmark for Vision-Centric Reasoning
Longteng Guo, Yifan Wang, Pengkang Huo +4
Recent multimodal large language models (MLLMs) achieve strong performance on visual reasoning benchmarks, yet it remains unclear to what extent such performance reflects reasoning…
Reason--Imagine--Act: Closed-Loop LLM Decision Making with World Models for Autonomous Driving
Zhengqi Sun, Yiwen Sun, Boxuan Liu +3
Large language models (LLMs) are promising for autonomous driving, but semantics-only decision policies can yield physically unsafe behavior in dynamic traffic. Existing methods ei…
Accelerating Streaming Video Large Language Models via Hierarchical Token Compression
Yiyu Wang, Xuyang Liu, Xiyan Gui +5
Streaming Video Large Language Models (VideoLLMs) have demonstrated impressive performance across various video understanding tasks, but they face significant challenges in real-ti…