2 papers
cs.CV2025
TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMs
Yunheng Li, Jing Cheng, Shaoyong Jia +4
This paper introduces TempSamp-R1, a new reinforcement fine-tuning framework designed to improve the effectiveness of adapting multimodal large language models (MLLMs) to video tem…
cs.CV2025
AR-1-to-3: Single Image to Consistent 3D Object Generation via Next-View Prediction
Xuying Zhang, Yupeng Zhou, Kai Wang +6
Novel view synthesis (NVS) is a cornerstone for image-to-3d creation. However, existing works still struggle to maintain consistency between the generated views and the input views…