activity
20242026
collaborators

9 papers

cs.LG2026

SALT: When More Rollouts Don't Help in Group-Based Policy Optimization and How to Make Them Matter

Powei Chang, Jinpeng Zhang, Chaoqun Sun +6

Reinforcement learning with verifiable rewards (RLVR) often adopts GRPO-style group-relative updates, sampling multiple rollouts per prompt to construct normalized learning signals…

cs.RO2026

Potential-Guided Flow Matching for Vision-Language-Action Policy Improvement

Yunpeng Mei, Jiakai He, Hongjie Cao +12

Large vision-language-action (VLA) policies are increasingly trained as conditional generative models over action chunks. Yet deployment produces mixed-quality experience-successfu…

cs.LG2026

SPICE: Submodular Penalized Information-Conflict Selection for Efficient Large Language Model Training

Powei Chang, Jinpeng Zhang, Bowen Chen +9

Information-based data selection for instruction tuning is compelling: maximizing the log-determinant of the Fisher information yields a monotone submodular objective, enabling gre…

cs.AI2025

SlimPack: Fine-Grained Asymmetric Packing for Balanced and Efficient Variable-Length LLM Training

Yuliang Liu, Guohao Wu, Shenglong Zhang +4

The efficient distributed training of Large Language Models (LLMs) is severely hampered by the extreme variance in context lengths. This data heterogeneity, amplified by convention…

cs.LG2025

Inpainting-Guided Policy Optimization for Diffusion Large Language Models

Siyan Zhao, Mengchen Liu, Jing Huang +8

Masked diffusion large language models (dLLMs) are emerging as promising alternatives to autoregressive LLMs, offering competitive performance while supporting unique generation ca…

cs.CV2025

FADE: Adversarial Concept Erasure in Flow Models

Zixuan Fu, Yan Ren, Finn Carter +5

Diffusion models have demonstrated remarkable image generation capabilities, but also pose risks in privacy and fairness by memorizing sensitive concepts or perpetuating biases. We…