5 papers
WebChallenger: A Reliable and Efficient Generalist Web Agent
Jayoo Hwang, Xiaowen Zhang, Vedant Padwal
Autonomous web navigation remains challenging for LLM agents, and the strongest generalist systems rely on proprietary reasoning models whose inference cost is prohibitive for the…
HyMem: Hybrid Memory Architecture with Dynamic Retrieval Scheduling
Xiaochen Zhao, Kaikai Wang, Xiaowen Zhang +2
Large language model (LLM) agents demonstrate strong performance in short-text contexts but often underperform in extended dialogues due to inefficient memory management. Existing…
AdaptMMBench: Benchmarking Adaptive Multimodal Reasoning for Mode Selection and Reasoning Process
Xintong Zhang, Xiaowen Zhang, Jingrong Wu +8
Adaptive multimodal reasoning has emerged as a promising frontier in Vision-Language Models (VLMs), aiming to dynamically modulate between tool-augmented visual reasoning and text…
Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs
Xintong Zhang, Zhi Gao, Bofei Zhang +8
Vision language models (VLMs) have achieved impressive performance across a variety of computer vision tasks. However, the multimodal reasoning capability has not been fully explor…
LightVA: Lightweight Visual Analytics with LLM Agent-Based Task Planning and Execution
Yuheng Zhao, Junjie Wang, Linbin Xiang +5
Visual analytics (VA) requires analysts to iteratively propose analysis tasks based on observations and execute tasks by creating visualizations and interactive exploration to gain…