activity
20232026
collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2026

SenseNova-MARS: Empowering Multimodal Agentic Reasoning and Search via Reinforcement Learning

Yong Xien Chng, Tao Hu, Wenwen Tong +10

While Vision-Language Models (VLMs) can solve complex tasks through agentic reasoning, their capabilities remain largely constrained to text-oriented chain-of-thought or isolated t…

cs.CV2025

On-the-fly Large-scale 3D Reconstruction from Multi-Camera Rigs

Yijia Guo, Tong Hu, Zhiwei Li +6

Recent advances in 3D Gaussian Splatting (3DGS) have enabled efficient free-viewpoint rendering and photorealistic scene reconstruction. While on-the-fly extensions of 3DGS have sh…

cs.CV2025

VideoSSM: Autoregressive Long Video Generation with Hybrid State-Space Memory

Yifei Yu, Xiaoshan Wu, Xinting Hu +8

Autoregressive (AR) diffusion enables streaming, interactive long-video generation by producing frames causally, yet maintaining coherence over minute-scale horizons remains challe…

cs.CV2025

A Unit Enhancement and Guidance Framework for Audio-Driven Avatar Video Generation

S. Z. Zhou, Y. B. Wang, J. F. Wu +2

Audio-driven human animation technology is widely used in human-computer interaction, and the emergence of diffusion models has further advanced its development. Currently, most me…

cs.CV2024

Unsupervised Cross-Domain Regression for Fine-grained 3D Game Character Reconstruction

Qi Wen, Xiang Wen, Hao Jiang +5

With the rise of the ``metaverse'' and the rapid development of games, it has become more and more critical to reconstruct characters in the virtual world faithfully. The immersive…

cs.CV2024

Towards Secure and Usable 3D Assets: A Novel Framework for Automatic Visible Watermarking

Gursimran Singh, Tianxi Hu, Mohammad Akbari +2

3D models, particularly AI-generated ones, have witnessed a recent surge across various industries such as entertainment. Hence, there is an alarming need to protect the intellectu…