activity
20242026
most citedHallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

5 citations · 5 across the 17 of their papers we have counts for

collaborators

19 papers

cs.AI2026

Beyond Coherence: Benchmarking Professional Editing-Technique Execution in Multi-Shot Audio-Video Generation

Tianyi Zeng, Junchao Liao, Yujie Wei +9

Recent multi-shot audio-video generators can produce increasingly coherent and cinematic outputs, but coherence does not imply the ability to execute editing techniques. Profession…

cs.RO2026

WAM-Diff2: Hierarchical AR-to-Diffusion Distillation for Highly Efficient Autonomous Driving VLA

Zhihao Zhu, Hanlin Shang, Mingwang Xu +6

Vision-Language-Action (VLA) models have emerged as a prominent paradigm for end-to-end autonomous driving; however, their efficient deployment is severely constrained by high comp…

cs.CV2026

GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling

Yixuan Lai, Tianjia Shao, Weijia Dou +2

Generating visually consistent multi-shot videos remains an open challenge. As videos span more shots, inconsistencies can accumulate across shots, causing entities that reappear a…

cs.CV2026

SlotMemory: Object-Centric KV Memory for Streaming Long-Video Generation

Weijia Dou, Hui Li, Jiahao Cui +3

Streaming video generation models typically rely on temporal-centric memory, which organizes historical context as raw frames, chunk segments, or unclustered tokens. This organizat…

eess.IV2026

ML-CLIPSim: Multi-Layer CLIP Similarity for Machine-Oriented Image Quality

Feng Ding, Haisheng Fu, Jie Liang +3

We study full-reference image quality assessment from a machine-centric perspective, where images are evaluated by how well they preserve information for downstream models. We form…

cs.CV2026

Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation

Chunyu Li, Jiaye Li, Ruiqiao Mei +4

Real-time text-driven joint audio-video avatar generation requires jointly synthesizing portrait video and speech with high fidelity and precise synchronization, yet existing audio…