activity
20212026
most citedCogView2: Faster and Better Text-to-Image Generation via Hierarchical Transformers

121 citations · 250 across the 12 of their papers we have counts for

collaborators
Showing cs.CVShow all

15 papers · 1 filter

cs.CV2026

Video2Code: Generating Interactive Webpages from UI Videos via Action-Aware Revisit

Mingde Xu, Zhen Yang, Yan Wang +7

UI videos provide a natural input for generating interactive webpages, as they capture both webpage appearance and action-triggered state transitions. However, directly applying vi…

cs.CV2025

UI2Code^N: UI-to-Code Generation as Interactive Visual Optimization

Zhen Yang, Wenyi Hong, Mingde Xu +5

UI-to-code aims to translate UI screenshots into executable front-end code. Despite progress with vision-language models (VLMs), most existing methods formulate UI-to-code as a sin…

cs.CV2025

Glyph: Scaling Context Windows via Visual-Text Compression

Jiale Cheng, Yusen Liu, Xinyu Zhang +11

Large language models (LLMs) increasingly rely on long-context modeling for tasks such as document understanding, code analysis, and multi-step reasoning. However, scaling context…

cs.CV2025

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models

Wenyi Hong, Yean Cheng, Zhuoyi Yang +6

In recent years, vision language models (VLMs) have made significant advancements in video understanding. However, a crucial capability - fine-grained motion comprehension - remain…

cs.CV20247 cited

CogVLM2: Visual Language Models for Image and Video Understanding

Wenyi Hong, Weihan Wang, Ming Ding +22

Beginning with VisualGLM and CogVLM, we are continuously exploring VLMs in pursuit of enhanced vision-language fusion, efficient higher-resolution architecture, and broader modalit…

cs.CV202418 cited

CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Zhuoyi Yang, Jiayan Teng, Wendi Zheng +15

We present CogVideoX, a large-scale text-to-video generation model based on diffusion transformer, which can generate 10-second continuous videos aligned with text prompt, with a f…