4 papers
Team-Based Self-Play With Dual Adaptive Weighting for Fine-Tuning LLMs
Wu Li, Yigeng Zhou, Zesheng Shi +3
While recent self-training approaches have reduced reliance on human-labeled data for aligning LLMs, they still face critical limitations: (i) sensitivity to synthetic data quality…
Offline Preference Optimization for Rectified Flow with Noise-Tracked Pairs
Yunhong Lu, Qichao Wang, Hengyuan Cao +2
Existing preference datasets for text-to-image models typically store only the final winner/loser images. This representation is insufficient for rectified flow (RF) models, whose…
Spherical Geometry Diffusion: Generating High-quality 3D Face Geometry via Sphere-anchored Representations
Junyi Zhang, Yiming Wang, Yunhong Lu +5
A fundamental challenge in text-to-3D face generation is achieving high-quality geometry. The core difficulty lies in the arbitrary and intricate distribution of vertices in 3D spa…
Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation
Yunhong Lu, Yanhong Zeng, Haobo Li +9
Efficient streaming video generation is critical for simulating interactive and dynamic worlds. Existing methods distill few-step video diffusion models with sliding window attenti…