52 citations · 142 across the 41 of their papers we have counts for
26 papers · 1 filter
RenderFormer-V2: Neural Rendering with Heterogeneous Scene Primitives
Chong Zeng, Yue Dong, Pieter Peers +2
We present 'RenderFormer-V2', a unified learned transformer-based neural rendering model, complementary to modern physics-based rendering systems, that can handle diverse light-tra…
Ring Forcing: Towards Precise Long-Term Memory for Autoregressive Video Diffusion
Bowen Xue, Brandon Y. Feng, Chenguo Lin +6
Scaling video generation to long durations reveals a critical bottleneck: current models lack robust long-term memory. This deficiency can be studied along two critical aspects: ob…
Masked Visual Actions for Unified World Modeling
Hadi Alzayer, Wenlong Huang, Haonan Chen +8
Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challe…
EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation
Songlin Yang, Haobin Zhong, Ruilin Zhang +23
The rapid evolution of generative video foundation models has propelled the field toward professional-grade cinematic synthesis. To achieve such demanding quality, the community tr…
Self-Consistency for LLM-Based Motion Trajectory Generation and Verification
Jiaju Ma, R. Kenny Jones, Jiajun Wu +1
Self-consistency has proven to be an effective technique for improving LLM performance on natural language reasoning tasks in a lightweight, unsupervised manner. In this work, we s…
Mode Seeking meets Mean Seeking for Fast Long Video Generation
Shengqu Cai, Weili Nie, Chao Liu +8
Scaling video generation from seconds to minutes faces a critical bottleneck: while short-video data is abundant and high-fidelity, coherent long-form data is scarce and limited to…