3 papers
cs.CV2026
TARS: Timestep-Aware Data Scaling for 3D-Free Video Re-Shooting
Jiwen Liu, Shujuan Li, Xiaohan Li +5
Video re-shooting aims to regenerate videos with controllable camera motion and viewpoint. Existing methods rely on explicit 3D priors, which are limited by reconstruction quality…
cs.CV2026
Vera: Identity-Faithful Human Subject-to-Video Generation
Yulong Xu, Xinyue Liu, Shujuan Li +6
Subject-to-video (S2V) generation has made substantial progress in preserving reference subjects across diverse categories, yet generic subject consistency remains insufficient for…
cs.CV2026
MVPBench: A Multi-Video Perception Evaluation Benchmark for Multi-Modal Video Understanding
Purui Bai, Tao Wu, Jiayang Sun +3
The rapid progress of Large Language Models (LLMs) has spurred growing interest in Multi-modal LLMs (MLLMs) and motivated the development of benchmarks to evaluate their perceptual…