7 papers
Prefix-Guided On-Policy Distillation: Mining Golden Trajectories from Rollouts
Qingfei Zhao, Huan Song, Shuyu Tian +2
On-policy distillation (OPD) improves reasoning models by applying dense teacher supervision on student-sampled trajectories. However, scaling OPD to long-horizon reasoning exposes…
BPPO: Binary Prefix Policy Optimization for Efficient GRPO-Style Reasoning RL with Concise Responses
Qingfei Zhao, Huan Song, Shuyu Tian +2
Group Relative Policy Optimization (GRPO) is widely used for training reasoning models, but updating all sampled completions in each group incurs substantial cost and can reinforce…
Ruyi2.5 Technical Report
Huan Song, Shuyu Tian, Qingfei Zhao +5
We present Ruyi2.5, a multimodal familial model built on the AI Flow framework. Extending Ruyi2's "Train Once, Deploy Many" paradigm to the multimodal domain, Ruyi2.5 constructs a…
Privacy-Aware Camera 2.0 Technical Report
Huan Song, Shuyu Tian, Ting Long +5
With the increasing deployment of intelligent sensing technologies in highly sensitive environments such as restrooms and locker rooms, visual surveillance systems face a profound…
Ruyi2 Technical Report
Huan Song, Shuyu Tian, Junyi Hao +5
Large Language Models (LLMs) face significant challenges regarding deployment costs and latency, necessitating adaptive computing strategies. Building upon the AI Flow framework, w…
A Real-Time Privacy-Preserving Behavior Recognition System via Edge-Cloud Collaboration
Huan Song, Shuyu Tian, Junyi Hao +4
As intelligent sensing expands into high-privacy environments such as restrooms and changing rooms, the field faces a critical privacy-security paradox. Traditional RGB surveillanc…