collaborators

7 papers

cs.CV2026

Vidu S1: A Real-Time Interactive Video Generation Model

Jintao Zhang, Kai Jiang, Jintao Chen +24

We introduce Vidu S1, a real-time interactive video generation model supporting voice control of digital characters. Users can control video generation content at any moment throug…

cs.CV2026

PhysiFormer: Learning to Simulate Mechanics in World Space

Yiming Chen, Yushi Lan, Andrea Vedaldi

We present PhysiFormer, a diffusion transformer for physically-plausible 3D object motion. Unlike video world models that operate in view-dependent pixel space, PhysiFormer represe…

cs.CV2026

SEED: Simple ViT and Evolving Harness for Explainable Text Forgery Detection

Kahim Wong, Kemou Li, Yiming Chen +2

AI-assisted image editing threatens trust in financial, legal, and identity records. The GenText-Forensics Challenge at ACM MM 2026 addresses this by requiring structured forensic…

cs.CV2026

minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models

Min Zhao, Hongzhou Zhu, Bokai Yan +9

Recent video diffusion foundation models have achieved remarkable progress in high-quality video generation, yet turning them into real-time interactive video world models remains…

cs.CV2025

Reasoning in Space via Grounding in the World

Yiming Chen, Zekun Qi, Wenyao Zhang +3

In this paper, we claim that 3D visual grounding is the cornerstone of spatial reasoning and introduce the Grounded-Spatial Reasoner (GS-Reasoner) to explore the effective spatial…

cs.CV2025

VicaSplat: A Single Run is All You Need for 3D Gaussian Splatting and Camera Estimation from Unposed Video Frames

Zhiqi Li, Chengrui Dong, Yiming Chen +2

We present VicaSplat, a novel framework for joint 3D Gaussians reconstruction and camera pose estimation from a sequence of unposed video frames, which is a critical yet underexplo…