collaborators

5 papers

cs.AI2026

FLY-EVAL++: An Evidence-Driven Evaluation Protocol for Safety-Constrained Flight Prediction with Large Language Models

Yalun Wu, Junfeng Fang, Jiawei Wang +6

Evaluating large language models (LLMs) in safety-critical, physics-governed environments requires more than accuracy-based metrics, because predictions that are numerically close…

cs.CV2026

CutVerse: A Compositional GUI Agents Benchmark for Media Post-Production Editing

Haobo Hu, Xiangwu Guo, Zhiheng Chen +4

While GUI agents have made significant progress in web navigation and basic operating system tasks, their capabilities in professional creative workflows remain largely underexplor…

cs.AI2026

PilotBench: A Benchmark for General Aviation Agents with Safety Constraints

Yalun Wu, Haotian Liu, Zhoujun Li +1

As Large Language Models (LLMs) advance toward embodied AI agents operating in physical environments, a fundamental question emerges: can models trained on text corpora reliably re…

cs.AI2026

SayNext-Bench: Why Do LLMs Struggle with Next-Utterance Anticipation?

Yueyi Yang, Haotian Liu, Fang Kang +4

We explore the use of large language models (LLMs) for next-utterance anticipation in human dialogue. Despite recent advances in LLMs demonstrating their ability to engage in natur…

eess.SY2025

Large Language Model as An Operator: An Experience-Driven Solution for Distribution Network Voltage Control

Xu Yang, Chenhui Lin, Licheng Sha +5

With the advanced reasoning, contextual understanding, and information synthesis capabilities of large language models (LLMs), a novel paradigm emerges for the autonomous generatio…