4 papers
SkillGrad: Optimizing Agent Skills Like Gradient Descent
Hanyu Wang, Yifan Lan, Bochuan Cao +2
Agent skills provide a lightweight way to adapt LLM agents to specialized domains by storing reusable procedural knowledge in structured files. However, whether downloaded from thi…
The Illusion of Reasoning: Exposing Evasive Data Contamination in LLMs via Zero-CoT Truncation
Yifan Lan, Yuanpu Cao, Hanyu Wang +2
Large language models (LLMs) have demonstrated impressive reasoning abilities across a wide range of tasks, but data contamination undermines the objective evaluation of these capa…
PreFlect: From Retrospective to Prospective Reflection in Large Language Model Agents
Hanyu Wang, Yuanpu Cao, Lu Lin +1
Advanced large language model agents typically adopt self-reflection for improving performance, where agents iteratively analyze past actions to correct errors. However, existing r…
TruthFlow: Truthful LLM Generation via Representation Flow Correction
Hanyu Wang, Bochuan Cao, Yuanpu Cao +1
Large language models (LLMs) are known to struggle with consistently generating truthful responses. While various representation intervention techniques have been proposed, these m…