collaborators

11 papers

cs.LG2026

How Much Dense Attention is Necessary? Oracle-Guided Sparse Prefill for Full/GQA Layers in Hybrid Long-Context Models

Hongxing Wang, Harenome Razanajato, Zhen Zhang +2

Long-context prefill remains expensive because full/GQA layers still score the historical sequence, even in hybrid models with local, sparse, linear, or recurrent components. We st…

cs.CV2026

LVSA: Training-Free Sparse Attention for Long Video Diffusion

Gael Glorian, Ioannis Lamprou, Zhen Zhang +2

Dense self-attention is the compute and quality bottleneck of long-video diffusion inference: cost grows quadratically with the sequence length, and beyond the training horizon the…

cs.DC2026

vLLM-Omni: Fully Disaggregated Serving for Any-to-Any Multimodal Models

Peiqi Yin, Jiangyun Zhu, Han Gao +13

Any-to-any multimodal models that jointly handle text, images, video, and audio represent a significant advance in multimodal AI. However, their complex architectures (typically co…

math.NA2025

PDEformer-2: A Versatile Foundation Model for Two-Dimensional Partial Differential Equations

Zhanhong Ye, Zining Liu, Bingyang Wu +7

Partial differential equations (PDEs) play a central role in describing many physical phenomena. Various scientific and engineering applications demand a versatile and differentiab…

cs.LG2025

Learnable-Differentiable Finite Volume Solver for Accelerated Simulation of Flows

Mengtao Yan, Qi Wang, Haining Wang +7

Simulation of fluid flows is crucial for modeling physical phenomena like meteorology, aerodynamics, and biomedicine. Classical numerical solvers often require fine spatiotemporal…

cs.CV2025

SlotPi: Physics-informed Object-centric Reasoning Models

Jian Li, Wan Han, Ning Lin +8

Understanding and reasoning about dynamics governed by physical laws through visual observation, akin to human capabilities in the real world, poses significant challenges. Current…