activity
20232026
most citedHSR-Enhanced Sparse Attention Acceleration

2 citations · 5 across the 22 of their papers we have counts for

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2026

UrbanAlign: Post-hoc Semantic Calibration for VLM-Human Preference Alignment

Yecheng Zhang, Rong Zhao, Zhizhou Sha +10

Vision-language models (VLMs) can describe urban scenes in rich detail, yet consistently fail to produce reliable human preference labels in domain-specific tasks such as safety as…

cs.CV2025

Discriminator-Free Direct Preference Optimization for Video Diffusion

Haoran Cheng, Qide Dong, Liang Peng +7

Direct Preference Optimization (DPO), which aligns models with human preferences through win/lose data pairs, has achieved remarkable success in language and image generation. Howe…

cs.CV2025

HOFAR: High-Order Augmentation of Flow Autoregressive Transformers

Yingyu Liang, Zhizhou Sha, Zhenmei Shi +2

Flow Matching and Transformer architectures have demonstrated remarkable performance in image generation tasks, with recent work FlowAR [Ren et al., 2024] synergistically integrati…

cs.CV2025

High-Order Matching for One-Step Shortcut Diffusion Models

Bo Chen, Chengyue Gong, Xiaoyu Li +5

One-step shortcut diffusion models [Frans, Hafner, Levine and Abbeel, ICLR 2025] have shown potential in vision generation, but their reliance on first-order trajectory supervision…

cs.CV2025

RichSpace: Enriching Text-to-Video Prompt Space via Text Embedding Interpolation

Yuefan Cao, Chengyue Gong, Xiaoyu Li +4

Text-to-video generation models have made impressive progress, but they still struggle with generating videos with complex features. This limitation often arises from the inability…

cs.CV2024

OmniControlNet: Dual-stage Integration for Conditional Image Generation

Yilin Wang, Haiyang Xu, Xiang Zhang +4

We provide a two-way integration for the widely adopted ControlNet by integrating external condition generation algorithms into a single dense prediction method and incorporating i…