21 papers
Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability
Shizhan Liu, Xinran Deng, Zhuoyi Yang +3
Latent diffusion models pair VAEs with diffusion backbones, and the structure of VAE latents strongly influences the difficulty of diffusion training. However, existing video VAEs…
Video2Code: Generating Interactive Webpages from UI Videos via Action-Aware Revisit
Mingde Xu, Zhen Yang, Yan Wang +7
UI videos provide a natural input for generating interactive webpages, as they capture both webpage appearance and action-triggered state transitions. However, directly applying vi…
UI2Code^N: UI-to-Code Generation as Interactive Visual Optimization
Zhen Yang, Wenyi Hong, Mingde Xu +5
UI-to-code aims to translate UI screenshots into executable front-end code. Despite progress with vision-language models (VLMs), most existing methods formulate UI-to-code as a sin…
GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents
V Team, Wenyi Hong, Xiaotao Gu +94
We present GLM-5V-Turbo, a step toward native foundation models for multimodal agents. As foundation models are increasingly deployed in real environments, agentic capability depen…
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models
Wenyi Hong, Yean Cheng, Zhuoyi Yang +6
In recent years, vision language models (VLMs) have made significant advancements in video understanding. However, a crucial capability - fine-grained motion comprehension - remain…
HoWToBench: Holistic Evaluation for LLM's Capability in Human-level Writing using Tree of Writing
Andrew Zhuoer Feng, Cunxiang Wang, Yu Luo +7
Evaluating the writing capabilities of large language models (LLMs) remains a significant challenge due to the multidimensional nature of writing skills and the limitations of exis…