3 papers
cs.AI2026
Benchmark Everything Everywhere All at Once
Shiyun Xiong, Dongming Wu, Peiwen Sun +5
Benchmarks are fundamental for evaluating and advancing LLMs and MLLMs by providing standardized and explicit measures of performance. However, their construction is labor-intensiv…
cs.CL2025
BehaviorSFT: Behavioral Token Conditioning for Clinical Agents Across the Proactivity Spectrum
Yubin Kim, Zhiyuan Hu, Hyewon Jeong +11
Large Language Models (LLMs) as clinical agents require careful behavioral adaptation. While adept at reactive tasks (e.g., diagnosis reasoning), LLMs often struggle with proactive…
cs.CL2025
Guiding VLM Agents with Process Rewards at Inference Time for GUI Navigation
Zhiyuan Hu, Shiyun Xiong, Yifan Zhang +5
Recent advancements in visual language models (VLMs) have notably enhanced their capabilities in handling complex Graphical User Interface (GUI) interaction tasks. Despite these im…