activity
20242026
collaborators

5 papers

cs.CV2026

SenseNova-MARS: Empowering Multimodal Agentic Reasoning and Search via Reinforcement Learning

Yong Xien Chng, Tao Hu, Wenwen Tong +10

While Vision-Language Models (VLMs) can solve complex tasks through agentic reasoning, their capabilities remain largely constrained to text-oriented chain-of-thought or isolated t…

cs.CV2025

On-the-fly Large-scale 3D Reconstruction from Multi-Camera Rigs

Yijia Guo, Tong Hu, Zhiwei Li +6

Recent advances in 3D Gaussian Splatting (3DGS) have enabled efficient free-viewpoint rendering and photorealistic scene reconstruction. While on-the-fly extensions of 3DGS have sh…

cs.CV2025

VideoSSM: Autoregressive Long Video Generation with Hybrid State-Space Memory

Yifei Yu, Xiaoshan Wu, Xinting Hu +8

Autoregressive (AR) diffusion enables streaming, interactive long-video generation by producing frames causally, yet maintaining coherence over minute-scale horizons remains challe…

cs.CV2025

A Unit Enhancement and Guidance Framework for Audio-Driven Avatar Video Generation

S. Z. Zhou, Y. B. Wang, J. F. Wu +2

Audio-driven human animation technology is widely used in human-computer interaction, and the emergence of diffusion models has further advanced its development. Currently, most me…

cs.CV2024

Unsupervised Cross-Domain Regression for Fine-grained 3D Game Character Reconstruction

Qi Wen, Xiang Wen, Hao Jiang +5

With the rise of the ``metaverse'' and the rapid development of games, it has become more and more critical to reconstruct characters in the virtual world faithfully. The immersive…