collaborators

23 papers

cs.AI2026

LinearKV: One Cached State Suffices for Position-Independent Caching in Hybrid LLMs

Yirui Liu, Ruoling Qi, Longwen Wang +5

LLM serving is increasingly accelerated by position-independent caching (PIC). Existing PIC methods, however, are built for full-attention models, where a token-indexed KV cache un…

cs.LG2026

Prefix-Guided On-Policy Distillation: Mining Golden Trajectories from Rollouts

Qingfei Zhao, Huan Song, Shuyu Tian +2

On-policy distillation (OPD) improves reasoning models by applying dense teacher supervision on student-sampled trajectories. However, scaling OPD to long-horizon reasoning exposes…

cs.LG2026

BPPO: Binary Prefix Policy Optimization for Efficient GRPO-Style Reasoning RL with Concise Responses

Qingfei Zhao, Huan Song, Shuyu Tian +2

Group Relative Policy Optimization (GRPO) is widely used for training reasoning models, but updating all sampled completions in each group incurs substantial cost and can reinforce…

cs.AI2026

Silicon Bureaucracy and AI Test-Oriented Education: Contamination Sensitivity and Score Confidence in LLM Benchmarks

Yiliang Song, Hongjun An, Jiangan Chen +4

Public benchmarks increasingly govern how large language models (LLMs) are ranked, selected, and deployed. We frame this benchmark-centered regime as Silicon Bureaucracy and AI Tes…

eess.IV2026

Enhancing Neural Video Compression of Static Scenes with Positive-Incentive Noise

Cheng Yuan, Zhenyu Jia, Jiawei Shao +1

Static scene videos, such as surveillance feeds and videotelephony streams, constitute a dominant share of storage consumption and network traffic. However, both traditional standa…

cs.CL2026

Ruyi2.5 Technical Report

Huan Song, Shuyu Tian, Qingfei Zhao +5

We present Ruyi2.5, a multimodal familial model built on the AI Flow framework. Extending Ruyi2's "Train Once, Deploy Many" paradigm to the multimodal domain, Ruyi2.5 constructs a…