Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Knowing How to Edit: Reliable Evaluation Signals for Diagnosing and Optimizing Prompts at Query Level
Ke Chen, Yifeng Wang, Hassan Almosapeeh +1
Prompt optimization has become a central mechanism for eliciting strong performance from LLMs, and recent work has made substantial progress by proposing diverse prompt evaluation…
cs.AI2025
Large Language Model-based Data Science Agent: A Survey
Ke Chen, Peiran Wang, Yaoning Yu +2
The rapid advancement of Large Language Models (LLMs) has driven novel applications across diverse domains, with LLM-based agents emerging as a crucial area of exploration. This su…
cs.AI2025
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems
Ke Chen, Yufei Zhou, Xitong Zhang +1
Automatic prompt generation plays a crucial role in enabling general-purpose multi-agent systems to perform diverse tasks autonomously. Existing methods typically evaluate prompts…