collaborators

11 papers

cs.SD2026

MMAE: A Massive Multitask Audio Editing Benchmark

Ziyang Ma, Ruiqi Yan, Ruiyang Xu +35

We introduce MMAE, a Massive Multitask Audio Editing benchmark, serving as the first comprehensive evaluation testbed designed for general-purpose instruction-based audio editing.…

cs.CL2026

EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle

Rong Wu, Xiaoman Wang, Jianbiao Mei +8

Current Large Language Model (LLM) agents show strong performance in tool use, but lack the crucial capability to systematically learn from their own experiences. While existing fr…

cs.IR2026

Bidirectional Semantic Complementary Tool Retrieval for Remote Sensing Agents

Zeyuan Wang, Dongyang Hou, Cheng Yang +10

Large language model (LLM)-based agents provide a novel paradigm for the automated processing of remote sensing(RS) data. Their success in complex RS tasks rely on extensive specia…

cs.AI2026

Mitigating Distribution Sharpening in Math RLVR via Distribution-Aligned Hint Synthesis and Backward Hint Annealing

Pei-Xi Xie, Che-Yu Lin, Cheng-Lin Yang

Reinforcement learning with verifiable rewards (RLVR) can improve low- reasoning accuracy while narrowing solution coverage on challenging math questions, and pass@1 gains do no…

cs.CR2026

TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories

Yen-Shan Chen, Sian-Yao Huang, Cheng-Lin Yang +1

As large language models (LLMs) evolve from static chatbots into autonomous agents, the primary vulnerability surface shifts from final outputs to intermediate execution traces. Wh…

cs.LG2026

GNNVerifier: Graph-based Verifier for LLM Task Planning

Yu Hao, Qiuyu Wang, Cheng Yang +3

Large language models (LLMs) facilitate the development of autonomous agents. As a core component of such agents, task planning aims to decompose complex natural language requests…