7 papers
TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming
Yibo Hu, Yu Qian, Mao Gu +6
E-commerce live streaming requires omni-modal understanding of noisy, temporally extended streams, where product facts are distributed across speech, video frames, product images,…
TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation
Qijun Gan, Chenwei Zhang, Meiguang Jin +2
Real-time long-form digital-human generation relies on causal models to extend audio-visual content while preserving subject appearance and audio-video synchronization across succe…
AptAvatar: Fast and Vivid Long-Form Audio-Driven Video Generation for Production-Ready Avatars
Hengyuan Zhang, Jingna Sun, Meiguang Jin +1
Production-ready audio-driven avatar generation requires efficient inference without sacrificing fidelity or motion expressiveness. However, existing acceleration methods often com…
OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models
Jinsen Su, Yongdong Luo, Yuexiao Ma +4
Existing token compression methods for omnimodal large language models typically rely on one modality to determine what to retain in the other. We show that this assumption often b…
Attention Grounded Enhancement for Visual Document Retrieval
Wanqing Cui, Wei Huang, Yazhi Guo +4
Visual document retrieval requires understanding heterogeneous and multi-modal content to satisfy implicit information needs. Recent advances use screenshot-based document encoding…
CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation
Xiangyang Luo, Xiaozhe Xin, Tao Feng +3
Synthesizing human--object interaction (HOI) videos has broad practical value in e-commerce, digital advertising, and virtual marketing. However, current diffusion models, despite…