activity
20242026
collaborators

18 papers

cs.CV2026

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation

Jiayu Chen, Xiaoyu Wu, Rongshan Gao +6

Audio-driven video generation (A2V) has achieved promising progress in synthesizing temporally coherent and audio-visually aligned videos, yet its inference remains expensive due t…

cs.CV2026

When and Where to Look: Adaptive Visual Evidence Scheduling for Efficient Long Video Understanding

Ke Li, Jiayu Chen, Maoliang Li +5

Efficient long-video understanding requires vision--language models (VLMs) to reason over a small number of frames selected as sparse visual evidence. Existing relevance-based meth…

cs.CV2026

Token Radius Attention for Efficient Video Generation

Jiayu Chen, Zhikun Jiang, Maoliang Li +6

Video Diffusion Transformers (VDiTs) enable high-fidelity generation but incur quadratic cost from dense 3D self-attention. Existing head- and block-level sparse methods share comp…

cs.CV2026

EcoVideo: Entropy-Orchestrated Video Generation Paradigm in Cloud-Edge Dynamics

Jiayu Chen, Hengyi Zhang, Maoliang Li +5

DiT video generation is latency-intensive due to iterative full-frame denoising, while prior cloud-edge methods largely rely on static inter-step decoupling and cannot leverage int…

cs.RO2026

VLN-Cache: Enabling Token Caching for VLN Models with Visual/Semantic Dynamics Awareness

Zihao Zheng, Zhihao Mao, Xingyue Zhou +9

Vision-and-Language Navigation (VLN) increasingly relies on large vision-language models, but their inference cost conflicts with real-time deployment. Token caching is a promising…

cs.DC2026

RoboECC: Multi-Factor-Aware Edge-Cloud Collaborative Deployment for VLA Models

Zihao Zheng, Hangyu Cao, Jiayu Chen +6

Vision-Language-Action (VLA) models are mainstream in embodied intelligence but face high inference costs. Edge-Cloud Collaborative (ECC) deployment offers an effective fix by easi…