Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
VP-VAE: Rethinking Vector Quantization via Adaptive Vector Perturbation
Linwei Zhai, Han Ding, Mingzhi Lin +5
Vector Quantized Variational Autoencoders (VQ-VAEs) are fundamental to modern generative modeling, yet they often suffer from training instability and "codebook collapse" due to th…
cs.LG2026
DisCO: Reinforcing Large Reasoning Models with Discriminative Constrained Optimization
Gang Li, Ming Lin, Tomer Galanti +2
The recent success and openness of DeepSeek-R1 have brought widespread attention to Group Relative Policy Optimization (GRPO) as a reinforcement learning method for large reasoning…
cs.LG2025
Model Steering: Learning with a Reference Model Improves Generalization Bounds and Scaling Laws
Xiyuan Wei, Ming Lin, Fanjiang Ye +4
This paper formalizes an emerging learning paradigm that uses a trained model as a reference to guide and enhance the training of a target model through strategic data selection or…