3 papers
cs.CL2025
Can LLMs Correct Themselves? A Benchmark of Self-Correction in LLMs
Guiyao Tie, Zenghui Yuan, Zeli Zhao +11
Self-correction of large language models (LLMs) emerges as a critical component for enhancing their reasoning performance. Although various self-correction methods have been propos…
cs.AI2025
SafeAgent: Safeguarding LLM Agents via an Automated Risk Simulator
Xueyang Zhou, Weidong Wang, Lin Lu +7
Large Language Model (LLM)-based agents are increasingly deployed in real-world applications such as "digital assistants, autonomous customer service, and decision-support systems"…
cs.CL2024
Way to Specialist: Closing Loop Between Specialized LLM and Evolving Domain Knowledge Graph
Yutong Zhang, Lixing Chen, Shenghong Li +6
Large language models (LLMs) have demonstrated exceptional performance across a wide variety of domains. Nonetheless, generalist LLMs continue to fall short in reasoning tasks nece…