2 citations · 5 across the 22 of their papers we have counts for
8 papers · 1 filter
UrbanAlign: Post-hoc Semantic Calibration for VLM-Human Preference Alignment
Yecheng Zhang, Rong Zhao, Zhizhou Sha +10
Vision-language models (VLMs) can describe urban scenes in rich detail, yet consistently fail to produce reliable human preference labels in domain-specific tasks such as safety as…
Discriminator-Free Direct Preference Optimization for Video Diffusion
Haoran Cheng, Qide Dong, Liang Peng +7
Direct Preference Optimization (DPO), which aligns models with human preferences through win/lose data pairs, has achieved remarkable success in language and image generation. Howe…
HOFAR: High-Order Augmentation of Flow Autoregressive Transformers
Yingyu Liang, Zhizhou Sha, Zhenmei Shi +2
Flow Matching and Transformer architectures have demonstrated remarkable performance in image generation tasks, with recent work FlowAR [Ren et al., 2024] synergistically integrati…
High-Order Matching for One-Step Shortcut Diffusion Models
Bo Chen, Chengyue Gong, Xiaoyu Li +5
One-step shortcut diffusion models [Frans, Hafner, Levine and Abbeel, ICLR 2025] have shown potential in vision generation, but their reliance on first-order trajectory supervision…
RichSpace: Enriching Text-to-Video Prompt Space via Text Embedding Interpolation
Yuefan Cao, Chengyue Gong, Xiaoyu Li +4
Text-to-video generation models have made impressive progress, but they still struggle with generating videos with complex features. This limitation often arises from the inability…
OmniControlNet: Dual-stage Integration for Conditional Image Generation
Yilin Wang, Haiyang Xu, Xiang Zhang +4
We provide a two-way integration for the widely adopted ControlNet by integrating external condition generation algorithms into a single dense prediction method and incorporating i…