collaborators

11 papers

cs.CV2026

Surface Keypoint Representation for Multi-Object and Articulated Human-Object Interaction Generation

Xiaogang Peng, Zeyu Han, Zichong Meng +4

Daily activities require humans to coordinate whole-body motion with the motion of surrounding objects. Despite recent progress in human-object interaction (HOI) generation, most e…

cs.CV2026

Streaming Video Generation with Streaming Force Control

Hanhui Wang, Yiming Xie, Haiwen Feng +3

We introduce StreamForce, a streaming video generation framework that enables physically grounded control through continuous force inputs. Unlike prior video models that train sepa…

cs.CV2026

LASER: Layer-wise Scale Alignment for Training-Free Streaming 4D Reconstruction

Tianye Ding, Yiming Xie, Yiqing Liang +3

Recent feed-forward reconstruction models like VGGT and achieve impressive reconstruction quality but cannot process streaming videos due to quadratic memory complexity, lim…

cs.CV2026

SocialFusion: Addressing Social Degradation in Pre-trained Vision-Language Models

Hamza Tahboub, Weiyan Shi, Gang Hua +1

Understanding social interactions from visual cues is a fundamental challenge for a socially competent AI. While powerful pre-trained vision-language models (VLMs) have shown remar…

cs.CV2025

HouseCrafter: Lifting Floorplans to 3D Scenes with 2D Diffusion Model

Hieu T. Nguyen, Yiwen Chen, Vikram Voleti +2

We introduce HouseCrafter, a novel approach that can lift a floorplan into a complete large 3D indoor scene (e.g., a house). Our key insight is to adapt a 2D diffusion model, which…

cs.CL2025

XFacta: Contemporary, Real-World Dataset and Evaluation for Multimodal Misinformation Detection with Multimodal LLMs

Yuzhuo Xiao, Zeyu Han, Yuhan Wang +1

The rapid spread of multimodal misinformation on social media calls for more effective and robust detection methods. Recent advances leveraging multimodal large language models (ML…