collaborators

7 papers

cs.LG2026

Prefix-Guided On-Policy Distillation: Mining Golden Trajectories from Rollouts

Qingfei Zhao, Huan Song, Shuyu Tian +2

On-policy distillation (OPD) improves reasoning models by applying dense teacher supervision on student-sampled trajectories. However, scaling OPD to long-horizon reasoning exposes…

cs.LG2026

BPPO: Binary Prefix Policy Optimization for Efficient GRPO-Style Reasoning RL with Concise Responses

Qingfei Zhao, Huan Song, Shuyu Tian +2

Group Relative Policy Optimization (GRPO) is widely used for training reasoning models, but updating all sampled completions in each group incurs substantial cost and can reinforce…

cs.CV2026

Privacy-Aware Camera 2.0 Technical Report

Huan Song, Shuyu Tian, Ting Long +5

With the increasing deployment of intelligent sensing technologies in highly sensitive environments such as restrooms and locker rooms, visual surveillance systems face a profound…

cs.CL2026

Ruyi2 Technical Report

Huan Song, Shuyu Tian, Junyi Hao +5

Large Language Models (LLMs) face significant challenges regarding deployment costs and latency, necessitating adaptive computing strategies. Building upon the AI Flow framework, w…

cs.CR2026

A Real-Time Privacy-Preserving Behavior Recognition System via Edge-Cloud Collaboration

Huan Song, Shuyu Tian, Junyi Hao +4

As intelligent sensing expands into high-privacy environments such as restrooms and changing rooms, the field faces a critical privacy-security paradox. Traditional RGB surveillanc…

cs.LG2026

Theoretical Foundations of Scaling Law in Familial Models

Huan Song, Qingfei Zhao, Ting Long +4

Neural scaling laws have become foundational for optimizing large language model (LLM) training, yet they typically assume a single dense model output. This limitation effectively…