most citedSCAIL: Towards Studio-Grade Character Animation via In-Context Learning of 3D-Consistent Pose Representations

1 citations · 1 across the 6 of their papers we have counts for

collaborators

7 papers

cs.CV2026

Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability

Shizhan Liu, Xinran Deng, Zhuoyi Yang +3

Latent diffusion models pair VAEs with diffusion backbones, and the structure of VAE latents strongly influences the difficulty of diffusion training. However, existing video VAEs…

cs.AI2026

LongWebBench: Evaluating Structural and Functional Webpage Generation in Long-Horizon Settings

Yi Zhao, Zhen Yang, Mengpan Chen +5

Recent vision-language models (VLMs) have shown promising progress in generating webpages from visual inputs, yet existing evaluations mainly focus on short, single-screen, and lar…

cs.CV2026

UI2Code^N: UI-to-Code Generation as Interactive Visual Optimization

Zhen Yang, Wenyi Hong, Mingde Xu +5

UI-to-code aims to translate UI screenshots into executable front-end code. Despite progress with vision-language models (VLMs), most existing methods formulate UI-to-code as a sin…

cs.SE2026

Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification

Zehai He, Wenyi Hong, Zhen Yang +4

Recent advances in large language models have improved the capabilities of coding agents, yet systematic evaluation of complex, end-to-end website development remains limited. To a…

cs.CV20261 cited

SCAIL: Towards Studio-Grade Character Animation via In-Context Learning of 3D-Consistent Pose Representations

Wenhao Yan, Sheng Ye, Zhuoyi Yang +6

Achieving controllable character animation that meets studio-grade standards remains challenging despite recent progress. Existing approaches can transfer motion from a driving vid…

cs.CV2026

Kaleido: Open-Sourced Multi-Subject Reference Video Generation Model

Zhenxing Zhang, Jiayan Teng, Zhuoyi Yang +6

We present Kaleido, a subject-to-video~(S2V) generation framework, which aims to synthesize subject-consistent videos conditioned on multiple reference images of target subjects. D…