11 papers
UI2App: Benchmarking Visual Interaction Inference in Executable Web Application Generation
Grace Man Chen, Litao Guo, Yifan Wu +5
Large language models (LLMs) have demonstrated growing competence in web page generation. However, existing text-driven approaches rely on complex prompts that impose substantial d…
No Place to Hide: Benchmarking Video Hallucination with Background-Controlled Pairs
Haojian Huang, Harold Haodong Chen, Meng Luo +6
We introduce VidPair-Halluc, a new benchmark for evaluating video hallucination in large video models (LVMs) under rigorous and controlled conditions. Unlike previous benchmarks th…
LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation
Yukang Chen, Luozhou Wang, Wei Huang +13
We present LongLive-2.0, an NVFP4-based parallel infrastructure throughout the full training and inference workflow of long video generation, addressing speed and memory bottleneck…
RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data
Harold Haodong Chen, Sirui Chen, Yingjie Xu +2
The scalability of robotic manipulation is fundamentally bottlenecked by the scarcity of task-aligned physical interaction data. While vision-language models (VLMs) and video gener…
Focusable Monocular Depth Estimation
Yuxin Du, Tao Lin, Zile Zhong +7
Monocular depth foundation models generalize well across scenes, yet they are typically optimized with uniform pixel-wise objectives that do not distinguish user-specified or task-…
RectifiedHR: Enable Efficient High-Resolution Synthesis via Energy Rectification
Zhen Yang, Guibao Shen, Minyang Li +5
Diffusion models have achieved remarkable progress across various visual generation tasks. However, their performance significantly declines when generating content at resolutions…