activity
20232025
most citedBeyond Linear Approximations: A Novel Pruning Approach for Attention Matrix

3 citations · 34 across the 47 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2025

Can You Count to Nine? A Human Evaluation Benchmark for Counting Limits in Modern Text-to-Video Models

Xuyang Guo, Zekai Huang, Jiayan Huo +4

Generative models have driven significant progress in a variety of AI tasks, including text-to-video generation, where models like Video LDM and Stable Video Diffusion can produce…

cs.CV2025

HOFAR: High-Order Augmentation of Flow Autoregressive Transformers

Yingyu Liang, Zhizhou Sha, Zhenmei Shi +2

Flow Matching and Transformer architectures have demonstrated remarkable performance in image generation tasks, with recent work FlowAR [Ren et al., 2024] synergistically integrati…

cs.CV2025

Text-to-Image Diffusion Models Cannot Count, and Prompt Refinement Cannot Help

Xuyang Guo, Jiayan Huo, Yingyu Liang +4

Generative modeling is widely regarded as one of the most essential problems in today's AI community, with text-to-image generation having gained unprecedented real-world impacts.…

cs.CV2025

High-Order Matching for One-Step Shortcut Diffusion Models

Bo Chen, Chengyue Gong, Xiaoyu Li +5

One-step shortcut diffusion models [Frans, Hafner, Levine and Abbeel, ICLR 2025] have shown potential in vision generation, but their reliance on first-order trajectory supervision…

cs.CV2025

RichSpace: Enriching Text-to-Video Prompt Space via Text Embedding Interpolation

Yuefan Cao, Chengyue Gong, Xiaoyu Li +4

Text-to-video generation models have made impressive progress, but they still struggle with generating videos with complex features. This limitation often arises from the inability…