works on

From the 1 of 15 linked papers with an AI index.

collaborators
Showing cs.AIShow all

6 papers · 1 filter

cs.AI2026

Matching Supervision to the Student's Learning Capacity: A Unified Framework for On-Policy Self-Distillation

Yongkang Yang, Zhezheng Hao, Hong Zhang +8

On-policy self-distillation (OPSD) improves the reasoning abilities of LLMs by internalizing privileged context into model parameters through self-distillation. Two recent research…

cs.AI2026

Echo: Learning from Experience Data via User-Driven Refinement

Hande Dong, Xiaoyun Liang, Jiarui Yu +15

Static "human data" faces inherent limitations: it is expensive to scale and bounded by the knowledge of its creators. Continuous learning from "experience data" - interactions bet…

cs.AI2026

Implicit Safety Alignment from Crowd Preferences

Qian Lin, Daniel S. Brown

Reinforcement Learning from Human Feedback (RLHF) can reveal implicit objectives such as safety considerations that go beyond task completion. In this work, we focus on the common…

cs.AI2026

ReCreate: Reasoning and Creating Domain Agents Driven by Experience

Zhezheng Hao, Hong Wang, Jian Luo +6

Large Language Model agents are reshaping the industrial landscape. However, most practical agents remain human-designed because tasks differ widely, making them labor-intensive to…

cs.AI2026

Scheduling Your LLM Reinforcement Learning with Reasoning Trees

Hong Wang, Zhezheng Hao, Jian Luo +6

Using Reinforcement Learning with Verifiable Rewards (RLVR) to optimize Large Language Models (LLMs) can be conceptualized as progressively editing a query's `Reasoning Tree'. This…

cs.AI2025

Unleashing the True Potential of LLMs: A Feedback-Triggered Self-Correction with Long-Term Multipath Decoding

Jipeng Li, Zeyu Gao, Yubin Qi +3

Large Language Models (LLMs) have achieved remarkable performance across diverse tasks, yet their susceptibility to generating incorrect content during inference remains a critical…