7 papers
Prefix-Guided On-Policy Distillation: Mining Golden Trajectories from Rollouts
Qingfei Zhao, Huan Song, Shuyu Tian +2
On-policy distillation (OPD) improves reasoning models by applying dense teacher supervision on student-sampled trajectories. However, scaling OPD to long-horizon reasoning exposes…
BPPO: Binary Prefix Policy Optimization for Efficient GRPO-Style Reasoning RL with Concise Responses
Qingfei Zhao, Huan Song, Shuyu Tian +2
Group Relative Policy Optimization (GRPO) is widely used for training reasoning models, but updating all sampled completions in each group incurs substantial cost and can reinforce…
Privacy-Aware Camera 2.0 Technical Report
Huan Song, Shuyu Tian, Ting Long +5
With the increasing deployment of intelligent sensing technologies in highly sensitive environments such as restrooms and locker rooms, visual surveillance systems face a profound…
Ruyi2 Technical Report
Huan Song, Shuyu Tian, Junyi Hao +5
Large Language Models (LLMs) face significant challenges regarding deployment costs and latency, necessitating adaptive computing strategies. Building upon the AI Flow framework, w…
A Real-Time Privacy-Preserving Behavior Recognition System via Edge-Cloud Collaboration
Huan Song, Shuyu Tian, Junyi Hao +4
As intelligent sensing expands into high-privacy environments such as restrooms and changing rooms, the field faces a critical privacy-security paradox. Traditional RGB surveillanc…
Theoretical Foundations of Scaling Law in Familial Models
Huan Song, Qingfei Zhao, Ting Long +4
Neural scaling laws have become foundational for optimizing large language model (LLM) training, yet they typically assume a single dense model output. This limitation effectively…