collaborators

7 papers

cs.AI2026

TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming

Yibo Hu, Yu Qian, Mao Gu +6

E-commerce live streaming requires omni-modal understanding of noisy, temporally extended streams, where product facts are distributed across speech, video frames, product images,…

cs.CV2026

TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation

Qijun Gan, Chenwei Zhang, Meiguang Jin +2

Real-time long-form digital-human generation relies on causal models to extend audio-visual content while preserving subject appearance and audio-video synchronization across succe…

cs.CV2026

AptAvatar: Fast and Vivid Long-Form Audio-Driven Video Generation for Production-Ready Avatars

Hengyuan Zhang, Jingna Sun, Meiguang Jin +1

Production-ready audio-driven avatar generation requires efficient inference without sacrificing fidelity or motion expressiveness. However, existing acceleration methods often com…

cs.CV2026

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models

Jinsen Su, Yongdong Luo, Yuexiao Ma +4

Existing token compression methods for omnimodal large language models typically rely on one modality to determine what to retain in the other. We show that this assumption often b…

cs.IR2026

Attention Grounded Enhancement for Visual Document Retrieval

Wanqing Cui, Wei Huang, Yazhi Guo +4

Visual document retrieval requires understanding heterogeneous and multi-modal content to satisfy implicit information needs. Recent advances use screenshot-based document encoding…

cs.CV2026

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation

Xiangyang Luo, Xiaozhe Xin, Tao Feng +3

Synthesizing human--object interaction (HOI) videos has broad practical value in e-commerce, digital advertising, and virtual marketing. However, current diffusion models, despite…