activity
20242026
most citedSeeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes

1 citations · 1 across the 5 of their papers we have counts for

collaborators
Showing cs.CVShow all

15 papers · 1 filter

cs.CV2026

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs

Sicheng Xu, Yu Deng, Shoukang Hu +5

Video diffusion models have significantly advanced portrait video generation, yet their high computational demands limit their use in interactive applications. This work presents a…

cs.CV2026

Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models

Dong Chen, Fangyun Wei, Ziyu Wan +18

We introduce Lens, a 3.8B-parameter T2I model that achieves performance competitive with, and in several cases surpassing, state-of-the-art models with more than 6B parameters acro…

cs.CV20261 cited

Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes

Zhiyuan Feng, Zhaolu Kang, Qijie Wang +16

Vision-language models (VLMs) are essential to Embodied AI, enabling robots to perceive, reason, and act in complex environments. They also serve as the foundation for the recent V…

cs.CV2025

VASA-3D: Lifelike Audio-Driven Gaussian Head Avatars from a Single Image

Sicheng Xu, Guojun Chen, Jiaolong Yang +4

We propose VASA-3D, an audio-driven, single-shot 3D head avatar generator. This research tackles two major challenges: capturing the subtle expression details present in real human…

cs.CV2025

Incorporating Pre-trained Diffusion Models in Solving the Schrödinger Bridge Problem

Zhicong Tang, Tiankai Hang, Shuyang Gu +2

This paper aims to unify Score-based Generative Models (SGMs), also known as Diffusion models, and the Schrödinger Bridge (SB) problem through three reparameterization techniques:…

cs.CV2025

Gaussian Variation Field Diffusion for High-fidelity Video-to-4D Synthesis

Bowen Zhang, Sicheng Xu, Chuxin Wang +4

In this paper, we present a novel framework for video-to-4D generation that creates high-quality dynamic 3D content from single video inputs. Direct 4D diffusion modeling is extrem…