24 papers
When and Where to Look: Adaptive Visual Evidence Scheduling for Efficient Long Video Understanding
Ke Li, Jiayu Chen, Maoliang Li +5
Efficient long-video understanding requires vision--language models (VLMs) to reason over a small number of frames selected as sparse visual evidence. Existing relevance-based meth…
Token Radius Attention for Efficient Video Generation
Jiayu Chen, Zhikun Jiang, Maoliang Li +6
Video Diffusion Transformers (VDiTs) enable high-fidelity generation but incur quadratic cost from dense 3D self-attention. Existing head- and block-level sparse methods share comp…
EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation
Jiayu Chen, Xiaoyu Wu, Rongshan Gao +6
Audio-driven video generation (A2V) has achieved promising progress in synthesizing temporally coherent and audio-visually aligned videos, yet its inference remains expensive due t…
EcoVideo: Entropy-Orchestrated Video Generation Paradigm in Cloud-Edge Dynamics
Jiayu Chen, Hengyi Zhang, Maoliang Li +5
DiT video generation is latency-intensive due to iterative full-frame denoising, while prior cloud-edge methods largely rely on static inter-step decoupling and cannot leverage int…
MoECa: Aligning Feature Reuse with Expert Decomposition in Diffusion Transformers
Maoliang Li, Haojing Chen, Jiayu Chen +4
Diffusion Transformers with Mixture-of-Experts (DiT-MoE) improve model capacity under sparse activation, but diffusion inference is still bottlenecked by redundant computation acro…
HeRo: Adaptive Orchestration of Agentic RAG on Heterogeneous Mobile SoC
Maoliang Li, Jiayu Chen, Zihao Zheng +5
With the increasing computational capability of mobile devices, deploying agentic retrieval-augmented generation (RAG) locally on heterogeneous System-on-Chips (SoCs) has become a…