5 papers
SenseNova-MARS: Empowering Multimodal Agentic Reasoning and Search via Reinforcement Learning
Yong Xien Chng, Tao Hu, Wenwen Tong +10
While Vision-Language Models (VLMs) can solve complex tasks through agentic reasoning, their capabilities remain largely constrained to text-oriented chain-of-thought or isolated t…
On-the-fly Large-scale 3D Reconstruction from Multi-Camera Rigs
Yijia Guo, Tong Hu, Zhiwei Li +6
Recent advances in 3D Gaussian Splatting (3DGS) have enabled efficient free-viewpoint rendering and photorealistic scene reconstruction. While on-the-fly extensions of 3DGS have sh…
VideoSSM: Autoregressive Long Video Generation with Hybrid State-Space Memory
Yifei Yu, Xiaoshan Wu, Xinting Hu +8
Autoregressive (AR) diffusion enables streaming, interactive long-video generation by producing frames causally, yet maintaining coherence over minute-scale horizons remains challe…
A Unit Enhancement and Guidance Framework for Audio-Driven Avatar Video Generation
S. Z. Zhou, Y. B. Wang, J. F. Wu +2
Audio-driven human animation technology is widely used in human-computer interaction, and the emergence of diffusion models has further advanced its development. Currently, most me…
Unsupervised Cross-Domain Regression for Fine-grained 3D Game Character Reconstruction
Qi Wen, Xiang Wen, Hao Jiang +5
With the rise of the ``metaverse'' and the rapid development of games, it has become more and more critical to reconstruct characters in the virtual world faithfully. The immersive…