collaborators

6 papers

cs.LG2025

GARDO: Reinforcing Diffusion Models without Reward Hacking

Haoran He, Yuxiao Ye, Jie Liu +7

Fine-tuning diffusion models via online reinforcement learning (RL) has shown great potential for enhancing text-to-image alignment. However, since precisely specifying a ground-tr…

cs.RO2025

UrbanVLA: A Vision-Language-Action Model for Urban Micromobility

Anqi Li, Zhiyong Wang, Jiazhao Zhang +5

Urban micromobility applications, such as delivery robots, demand reliable navigation across large-scale urban environments while following long-horizon route instructions. This ta…

cs.CV2025

Toward Robust Medical Fairness: Debiased Dual-Modal Alignment via Text-Guided Attribute-Disentangled Prompt Learning for Vision-Language Models

Yuexuan Xia, Benteng Ma, Jiang He +3

Ensuring fairness across demographic groups in medical diagnosis is essential for equitable healthcare, particularly under distribution shifts caused by variations in imaging equip…

cs.LG2025

Efficient Controllable Diffusion via Optimal Classifier Guidance

Owen Oertell, Shikun Sun, Yiding Chen +3

The controllable generation of diffusion models aims to steer the model to generate samples that optimize some given objective functions. It is desirable for a variety of applicati…

cs.LG2025

Towards More Efficient, Robust, Instance-adaptive, and Generalizable Sequential Decision making

Zhiyong Wang

The primary goal of my Ph.D. study is to develop provably efficient and practical algorithms for data-driven sequential decision-making under uncertainty. My work focuses on reinfo…

cs.CV2025

VLIPP: Towards Physically Plausible Video Generation with Vision and Language Informed Physical Prior

Xindi Yang, Baolu Li, Yiming Zhang +8

Video diffusion models (VDMs) have advanced significantly in recent years, enabling the generation of highly realistic videos and drawing the attention of the community in their po…