collaborators

8 papers

cs.LG2026

Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling

Xinmu Ge, Zizhuo Zhang, Yu Huang +9

On-policy distillation (OPD) has emerged as a promising post-training technique for enhancing LLM reasoning. It is commonly believed to enable the student model to distill knowledg…

cs.CL2026

FinReportBench: Measuring and Improving Institution-Grade Financial Report Generation

Yinghao Tang, Tan Zhenwei, Yiyao Wang +4

Large language models can produce fluent financial analysis, but fluency alone does not establish whether a report is suitable for institutional delivery. We introduce FinReportBen…

cs.AI2026

LiveEvalBench: Toward Open-World Evaluation for Web Generation

Yiyao Wang, Zhen Wen, Yinghao Tang +5

Large language models are increasingly capable of synthesizing executable frontend projects, yet existing benchmarks still treat web generation as a static evaluation problem. We a…

cs.LG2026

Focal Reward: Balanced Reinforcement Learning under Rubric-Based Rewards

Yu Huang, Zihua Zhao, Zhaoxin Huan +9

The open-ended generation in LLMs usually requires multi-dimensional rubrics to adequately assess quality and guide the improvement of reinforcement learning. However, a critical d…

cs.CY2026

LSSF: Safety Alignment for Large Language Models through Low-Rank Safety Subspace Fusion

Guanghao Zhou, Panjia Qiu, Cen Chen +4

The safety mechanisms of large language models (LLMs) exhibit notable fragility, as even fine-tuning on datasets without harmful content may still undermine their safety capabiliti…

cs.AI2025

Understanding and Mitigating Overrefusal in LLMs from an Unveiling Perspective of Safety Decision Boundary

Licheng Pan, Yongqi Tong, Xin Zhang +3

Large language models (LLMs) have demonstrated remarkable capabilities across a wide range of tasks, yet they often refuse to answer legitimate queries--a phenomenon known as overr…