8 papers
Scaling with Confidence: Calibrating Confidence of LLMs for Adaptive Test Time Scaling
Xuqing Yang, Yi Yuan, Shanzhe Lei +1
Training large language models (LLMs) with reinforcement learning (RL) has significantly advanced their performance on reasoning and question-answering tasks. However, prevailing R…
GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents
Xiongbin Wu, Zhihao Luo, Shanzhe Lei +7
Recently, vision-language model (VLM) agents have shown promising progress in open-world tasks, where successful task completion often requires multiple turns of visual perception…
Safactory: A Scalable Agentic Infrastructure for Training Trustworthy Autonomous Intelligence
Xinquan Chen, Zhenyun Yin, Shan He +38
As large models evolve from conversational assistants into autonomous agents, challenges increasingly arise from long-horizon decision making, tool use, and real environment intera…
Towards Trustworthy Report Generation: A Deep Research Agent with Progressive Confidence Estimation and Calibration
Yi Yuan, Xuhong Wang, Shanzhe Lei
As agent-based systems continue to evolve, deep research agents are capable of automatically generating research-style reports across diverse domains. While these agents promise to…
UniMark: Artificial Intelligence Generated Content Identification Toolkit
Meilin Li, Ji He, Yi Yu +5
The rapid proliferation of Artificial Intelligence Generated Content has precipitated a crisis of trust and urgent regulatory demands. However, existing identification tools suffer…
Beyond Correctness: Confidence-Aware Reward Modeling for Enhancing Large Language Model Reasoning
Qianxi He, Qingyu Ren, Shanzhe Lei +2
Recent advancements in large language models (LLMs) have shifted the post-training paradigm from traditional instruction tuning and human preference alignment toward reinforcement…