activity
20242026
most citedDo LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMs

3 citations · 3 across the 4 of their papers we have counts for

collaborators

8 papers

cs.LG2026

Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Siyan Zhao, Zhihui Xie, Mengchen Liu +4

Knowledge distillation improves large language model (LLM) reasoning by compressing the knowledge of a teacher LLM to train smaller LLMs. On-policy distillation advances this appro…

cs.CV2025

Accelerating Inference of Masked Image Generators via Reinforcement Learning

Pranav Subbaraman, Shufan Li, Siyan Zhao +1

Masked Generative Models (MGM)s demonstrate strong capabilities in generating high-fidelity images. However, they need many sampling steps to create high-quality generations, resul…

cs.CL2025

SPG: Sandwiched Policy Gradient for Masked Diffusion Language Models

Chenyu Wang, Paria Rashidinejad, DiJia Su +9

Diffusion large language models (dLLMs) are emerging as an efficient alternative to autoregressive models due to their ability to decode multiple tokens in parallel. However, align…

cs.LG2025

Inpainting-Guided Policy Optimization for Diffusion Large Language Models

Siyan Zhao, Mengchen Liu, Jing Huang +8

Masked diffusion large language models (dLLMs) are emerging as promising alternatives to autoregressive LLMs, offering competitive performance while supporting unique generation ca…

cs.LG2025

Multi-fidelity Reinforcement Learning Control for Complex Dynamical Systems

Luning Sun, Xin-Yang Liu, Siyan Zhao +3

Controlling instabilities in complex dynamical systems is challenging in scientific and engineering applications. Deep reinforcement learning (DRL) has seen promising results for a…

cs.CL2025

d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement Learning

Siyan Zhao, Devaansh Gupta, Qinqing Zheng +1

Recent large language models (LLMs) have demonstrated strong reasoning capabilities that benefits from online reinforcement learning (RL). These capabilities have primarily been de…