5 papers
CAD: Conflict-Aware Decoding to Mitigate Cross-Modal Hallucinations in Omnimodal Large Language Models
Yuchen Deng, Chang Sun, Hai-Tao Zheng +2
Omnimodal large language models (Omni-LLMs) integrate audio, video, and text, yet remain vulnerable to cross-modal hallucinations, where one modality improperly influences predicti…
GeoEdit: Geometry-Aware Object Editing via Dual-Branch Denoising
Yi He, Jiangming Wang, Xinyu Wang +5
Precisely manipulating objects in a single photograph (translation, rotation, scaling) while obeying 3D physical constraints remains unsolved for diffusion-based editors. Current 2…
OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models
Yuchen Deng, Zidang Cai, Hai-Tao Zheng +3
Omnimodal large language models (Omni-LLMs) show strong capability in audio-video understanding, but their practical deployment remains limited by high inference cost of long video…
Beyond Boundary Frames: Talking-Head Inbetweening via Context-Aware Motion Modeling
Yuchen Deng, Hai-Tao Zheng, Jie Wang +3
Existing talking-head generation methods primarily target open-ended generation rather than bridging two existing video segments. In this paper, we study talking-head inbetweening,…
FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling
Yuchen Deng, Xiuyang Wu, Hai-Tao Zheng +3
Current talking-head generation has gradually shifted from GAN-based methods to diffusion-based paradigms, achieving remarkable progress in visual fidelity and temporal consistency…