18 papers · 1 filter
SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models
Chenhao Dang, Siyuan Xiong, Conghui He +1
Agent skills represent a standardized format for packaging procedural knowledge and domain expertise, serving within agent harness systems as an essential mechanism to continually…
Reinforcement Learning for Evidence-Seeking Diagnostic Reasoning with Large Language Models
Shengyi Hua, Kangzhe Hu, Conghui He +2
Recent reasoning-centric Large Language Models (LLMs) have made significant strides, yet they predominantly operate on a passive-inference pattern that assumes complete information…
PAGER: Bridging the Semantic-Execution Gap in Point-Precise Geometric GUI Control
Jingxuan Wei, Xi Bai, Shan Liu +8
Large vision-language models have significantly advanced GUI agents, enabling executable interaction across web, mobile, and desktop interfaces. Yet these gains largely rely on a f…
NanoResearch: Co-Evolving Skills, Memory, and Policy for Personalized Research Automation
Jinhang Xu, Qiyuan Zhu, Yujun Wu +11
LLM-powered multi-agent systems can now automate the full research pipeline from ideation to paper writing, but a fundamental question remains: automation for whom? Researchers ope…
Respecting Self-Uncertainty in On-Policy Self-Distillation for Efficient LLM Reasoning
Junlong Ke, Zichen Wen, Weijia Li +2
On-policy self-distillation trains a reasoning model on its own rollouts while a teacher, often the same model conditioned on privileged context, provides dense token-level supervi…
PaperFit: Vision-in-the-Loop Typesetting Optimization for Scientific Documents
Bihui Yu, Xinglong Xu, Junjie Jiang +6
A LaTeX manuscript that compiles without error is not necessarily publication-ready. The resulting PDFs frequently suffer from misplaced floats, overflowing equations, inconsistent…