4 papers · 1 filter
Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses?
Jinyi Han, Yuanjian Xu, Ying Liao +6
Large language model (LLM) agents increasingly rely on skills, structured documents that specify when to act, which procedure to follow, and which tools are allowed. Existing evalu…
Terminal-World: Scaling Terminal-Agent Environments via Agent Skills
Zihao Cheng, Hongru Wang, Zeming Liu +6
Terminal agents extend Large Language Models with the ability to execute tasks directly in command-line environments, but their progress is bottlenecked by the scarcity of high-qua…
A Stitch in Time Saves Nine: Proactive Self-Refinement for Language Models
Jinyi Han, Xinyi Wang, Haiquan Zhao +9
Recent advances in self-refinement have demonstrated significant potential for improving the outputs of large language models (LLMs) through iterative refinement. However, most exi…
Mind the Generation Process: Fine-Grained Confidence Estimation During LLM Generation
Jinyi Han, Tingyun Li, Shisong Chen +8
While large language models (LLMs) have demonstrated remarkable performance across diverse tasks, they fundamentally lack self-awareness and frequently exhibit overconfidence, assi…