23 papers
BehaviorBench: Modeling Real-World User Decisions from Behavioral Traces
Liangwei Yang, Jielin Qiu, Zixiang Chen +9
Many decision-support settings require systems that adapt to individual users, but evaluation data for this problem remain limited. Existing benchmarks for user understanding often…
DIG to Heal: Scaling General-purpose Agent Collaboration via Explainable Dynamic Decision Paths
Hanqing Yang, Hyungwoo Lee, Yuhang Yao +4
The increasingly popular agentic AI paradigm promises to harness the power of multiple, general-purpose large language model (LLM) agents to collaboratively complete complex tasks.…
CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?
Haolin Chen, Deon Metelski, Leon Qi +30
End-to-end automation of realistic healthcare operations stresses three capabilities underrepresented in current benchmarks: policy density, decisions must be grounded in a large l…
Towards Scalable Lightweight GUI Agents via Multi-role Orchestration
Ziwei Wang, Junjie Zheng, Leyang Yang +7
Autonomous Graphical User Interface (GUI) agents powered by Multimodal Large Language Models (MLLMs) enable digital automation on end-user devices. While scaling both parameters an…
Whisper-AuT: Domain-Adapted Audio Encoder for Efficient Audio-LLM Training
Jielin Qiu, Ming Zhu, Wenting Zhao +11
Audio-native large language models (audio-LLMs) commonly use Whisper as their audio encoder. However, Whisper was trained exclusively on speech data, producing weak representations…
Webscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining Levels
Zhepeng Cen, Haolin Chen, Shiyu Wang +8
Large Language Models (LLMs) have achieved remarkable success through imitation learning on vast text corpora, but this paradigm creates a training-generation gap and limits robust…