4 papers
CutVerse: A Compositional GUI Agents Benchmark for Media Post-Production Editing
Haobo Hu, Xiangwu Guo, Zhiheng Chen +4
While GUI agents have made significant progress in web navigation and basic operating system tasks, their capabilities in professional creative workflows remain largely underexplor…
SayNext-Bench: Why Do LLMs Struggle with Next-Utterance Anticipation?
Yueyi Yang, Haotian Liu, Fang Kang +4
We explore the use of large language models (LLMs) for next-utterance anticipation in human dialogue. Despite recent advances in LLMs demonstrating their ability to engage in natur…
Large Language Model as An Operator: An Experience-Driven Solution for Distribution Network Voltage Control
Xu Yang, Chenhui Lin, Licheng Sha +5
With the advanced reasoning, contextual understanding, and information synthesis capabilities of large language models (LLMs), a novel paradigm emerges for the autonomous generatio…
PilotBench: A Benchmark for General Aviation Agents with Safety Constraints
Yalun Wu, Haotian Liu, Zhoujun Li +1
As Large Language Models (LLMs) advance toward embodied AI agents operating in physical environments, a fundamental question emerges: can models trained on text corpora reliably re…