6 papers
OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models
Yuchen Deng, Zidang Cai, Hai-Tao Zheng +3
Omnimodal large language models (Omni-LLMs) show strong capability in audio-video understanding, but their practical deployment remains limited by high inference cost of long video…
FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling
Yuchen Deng, Xiuyang Wu, Hai-Tao Zheng +3
Current talking-head generation has gradually shifted from GAN-based methods to diffusion-based paradigms, achieving remarkable progress in visual fidelity and temporal consistency…
Online3R: Online Learning for Consistent Sequential Reconstruction Based on Geometry Foundation Model
Shunkai Zhou, Zike Yan, Fei Xue +3
We present Online3R, a new sequential reconstruction framework that is capable of adapting to new scenes through online learning, effectively resolving inconsistency issues. Specif…
ProCap: Projection-Aware Captioning for Spatial Augmented Reality
Zimo Cao, Yuchen Deng, Haibin Ling +1
Spatial augmented reality (SAR) directly projects digital content onto physical scenes using projectors, creating immersive experience without head-mounted displays. However, for S…
Beyond Boundary Frames: Talking-Head Inbetweening via Context-Aware Motion Modeling
Yuchen Deng, Xiuyang Wu, Hai-Tao Zheng +4
Existing talking-head generation methods primarily target open-ended generation rather than bridging two existing video segments. In this paper, we study talking-head inbetweening,…
LAPIG: Language Guided Projector Image Generation with Surface Adaptation and Stylization
Yuchen Deng, Haibin Ling, Bingyao Huang
We propose LAPIG, a language guided projector image generation method with surface adaptation and stylization. LAPIG consists of a projector-camera system and a target textured pro…