collaborators

8 papers

cs.CV2026

When and Where to Look: Adaptive Visual Evidence Scheduling for Efficient Long Video Understanding

Ke Li, Jiayu Chen, Maoliang Li +5

Efficient long-video understanding requires vision--language models (VLMs) to reason over a small number of frames selected as sparse visual evidence. Existing relevance-based meth…

cs.CV2026

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation

Jiayu Chen, Xiaoyu Wu, Rongshan Gao +6

Audio-driven video generation (A2V) has achieved promising progress in synthesizing temporally coherent and audio-visually aligned videos, yet its inference remains expensive due t…

astro-ph.CO2026

DESI DR2 Baryon Acoustic Oscillations from the Lyman Alpha Forest Multipoles

N. G. Karaçaylı, A. Cuceu, J. Aguilar +55

We present an alternative measurement of the Baryon Acoustic Oscillation (BAO) using the Legendre multipole representation of the Ly forest correlation functions from the secon…

cs.LG2026

MoECa: Aligning Feature Reuse with Expert Decomposition in Diffusion Transformers

Maoliang Li, Haojing Chen, Jiayu Chen +4

Diffusion Transformers with Mixture-of-Experts (DiT-MoE) improve model capacity under sparse activation, but diffusion inference is still bottlenecked by redundant computation acro…

cs.RO2026

VLN-Cache: Enabling Token Caching for VLN Models with Visual/Semantic Dynamics Awareness

Zihao Zheng, Zhihao Mao, Xingyue Zhou +9

Vision-and-Language Navigation (VLN) increasingly relies on large vision-language models, but their inference cost conflicts with real-time deployment. Token caching is a promising…

cs.AR2026

GOMA: Geometrically Optimal Mapping via Analytical Modeling for Spatial Accelerators

Wulve Yang, Hailong Zou, Rui Zhou +5

General matrix multiplication (GEMM) on spatial accelerators is highly sensitive to mapping choices in both execution efficiency and energy consumption. However, the mapping space…