5 papers
RoPE Attention Can Be Trained in Almost Linear Time
Yang Cao, Jiayan Huo, Yingyu Liang +2
The Rotary Position Embedding (RoPE) mechanism has become a powerful enhancement to the Transformer architecture, which enables models to capture token relationships when encoding…
Text-to-Image Diffusion Models Cannot Count, and Prompt Refinement Cannot Help
Xuyang Guo, Jiayan Huo, Yingyu Liang +4
Generative modeling is widely regarded as one of the most essential problems in today's AI community, with text-to-image generation having gained unprecedented real-world impacts.…
T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models
Xuyang Guo, Jiayan Huo, Zhenmei Shi +3
Thanks to recent advancements in scalable deep architectures and large-scale pretraining, text-to-video generation has achieved unprecedented capabilities in producing high-fidelit…
T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation
Xuyang Guo, Jiayan Huo, Zhenmei Shi +3
Text-to-video generative models have made significant strides in recent years, producing high-quality videos that excel in both aesthetic appeal and accurate instruction following,…
Can You Count to Nine? A Human Evaluation Benchmark for Counting Limits in Modern Text-to-Video Models
Xuyang Guo, Zekai Huang, Jiayan Huo +4
Generative models have driven significant progress in a variety of AI tasks, including text-to-video generation, where models like Video LDM and Stable Video Diffusion can produce…