collaborators

6 papers

cs.AI2026

DRPO: Efficient Reasoning via Decoupled Reward Policy Optimization

Gang Li, Yan Chen, Ming Lin +1

Recent large reasoning models (LRMs) driven by reinforcement learning algorithms (e.g., GRPO) have achieved remarkable performance on challenging reasoning tasks. However, these mo…

cs.LG2026

VP-VAE: Rethinking Vector Quantization via Adaptive Vector Perturbation

Linwei Zhai, Han Ding, Mingzhi Lin +5

Vector Quantized Variational Autoencoders (VQ-VAEs) are fundamental to modern generative modeling, yet they often suffer from training instability and "codebook collapse" due to th…

cs.LG2026

DisCO: Reinforcing Large Reasoning Models with Discriminative Constrained Optimization

Gang Li, Ming Lin, Tomer Galanti +2

The recent success and openness of DeepSeek-R1 have brought widespread attention to Group Relative Policy Optimization (GRPO) as a reinforcement learning method for large reasoning…

cs.CL2025

Beyond the Black Box: Demystifying Multi-Turn LLM Reasoning with VISTA

Yiran Zhang, Mingyang Lin, Mark Dras +1

Recent research has increasingly focused on the reasoning capabilities of Large Language Models (LLMs) in multi-turn interactions, as these scenarios more closely mirror real-world…

cs.CY2025

PHORECAST: Enabling AI Understanding of Public Health Outreach Across Populations

Rifaa Qadri, Anh Nhat Nhu, Swati Ramnath +6

Understanding how diverse individuals and communities respond to persuasive messaging holds significant potential for advancing personalized and socially aware machine learning. Wh…

cs.LG2025

Model Steering: Learning with a Reference Model Improves Generalization Bounds and Scaling Laws

Xiyuan Wei, Ming Lin, Fanjiang Ye +4

This paper formalizes an emerging learning paradigm that uses a trained model as a reference to guide and enhance the training of a target model through strategic data selection or…