collaborators

17 papers

cs.CV2026

AdaForensics: Learning A Characteristic-aware Adaptive Deepfake Detector

Xiaoke Yang, Haixu Song, Xiangyu Lu +2

In this paper, we propose a characteristic-aware adaptive network named AdaForensics for deepfake detection. Most existing methods learn a fixed network to detect deepfakes based o…

cs.CV2026

Learning Efficient 4D Gaussian Representations from Monocular Videos with Flow Splatting

Shengjun Zhang, Jinzhao Li, Xin Fei +1

Reconstructing dynamic 3D scenes from monocular videos is challenging due to scene complexity and temporal dynamics. With the advancement of 3D Gaussian Splatting in novel view syn…

cs.CV2026

Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players

Fangfu Liu, Kai He, Tianchang Shen +7

World models for interactive video generation have largely focused on single-agent settings, where future observations are generated from a single control signal. However, many gen…

cs.CV2026

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence

Diankun Wu, Fangfu Liu, Yi-Hsin Hung +1

Recent advancements in Multimodal Large Language Models (MLLMs) have significantly enhanced performance on 2D visual tasks. However, improving their spatial intelligence remains a…

cs.CV2026

Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling

Haoyu Wu, Diankun Wu, Tianyu He +4

Videos inherently represent 2D projections of a dynamic 3D world. However, our analysis suggests that video diffusion models trained solely on raw video data often fail to capture…

cs.CV2026

Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training

Fangfu Liu, Diankun Wu, Jiawei Chi +7

Humans perceive and understand real-world spaces through a stream of visual observations. Therefore, the ability to streamingly maintain and update spatial evidence from potentiall…