collaborators

9 papers

cs.CV2026

Text-to-Image Diffusion Models Cannot Count, and Prompt Refinement Cannot Help

Xuyang Guo, Jiayan Huo, Yingyu Liang +4

Generative modeling is widely regarded as one of the most essential problems in today's AI community, with text-to-image generation having gained unprecedented real-world impacts.…

cs.AI2025

Evaluating Frontier LLMs on PhD-Level Mathematical Reasoning: A Benchmark on a Textbook in Theoretical Computer Science about Randomized Algorithms

Yang Cao, Yubin Chen, Xuyang Guo +4

The rapid advancement of large language models (LLMs) has led to significant breakthroughs in automated mathematical reasoning and scientific discovery. Georgiev, Gmez-Serran…

cs.CV2025

Your Vision-Language Model Can't Even Count to 20: Exposing the Failures of VLMs in Compositional Counting

Xuyang Guo, Zekai Huang, Zhenmei Shi +2

Vision-Language Models (VLMs) have become a central focus of today's AI community, owing to their impressive abilities gained from training on large-scale vision-language data from…

cs.CR2025

Too Easily Fooled? Prompt Injection Breaks LLMs on Frustratingly Simple Multiple-Choice Questions

Xuyang Guo, Zekai Huang, Zhao Song +1

Large Language Models (LLMs) have recently demonstrated strong emergent abilities in complex reasoning and zero-shot generalization, showing unprecedented potential for LLM-as-a-ju…

cs.CV2025

T2VWorldBench: A Benchmark for Evaluating World Knowledge in Text-to-Video Generation

Yubin Chen, Xuyang Guo, Zhenmei Shi +2

Text-to-video (T2V) models have shown remarkable performance in generating visually reasonable scenes, while their capability to leverage world knowledge for ensuring semantic cons…

cs.CV2025

T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models

Xuyang Guo, Jiayan Huo, Zhenmei Shi +3

Thanks to recent advancements in scalable deep architectures and large-scale pretraining, text-to-video generation has achieved unprecedented capabilities in producing high-fidelit…