7 papers
Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification
Haopeng Li, Yitong Li, Junsong Chen +8
Diffusion transformers are essential for high-fidelity video generation, but long token sequences make attention a dominant inference bottleneck. Training-free dynamic sparse atten…
Sol Video Inference Engine: Agent-Native Full-Stack Acceleration Framework for Efficient Video Generation
Yitong Li, Junsong Chen, Haopeng Li +6
Modern video diffusion models achieve higher generation quality through scaling, but this also increases inference cost. Although many acceleration methods have been proposed, a ce…
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model
Lichen Bai, Tianhao Zhang, Shitong Shao +14
As an increasing majority of global video content is consumed on social platforms for interactive social purposes, video generation models built for social worlds are important but…
Autoregressive Image Generation with Randomized Parallel Decoding
Haopeng Li, Jinyue Yang, Guoqi Li +1
We introduce ARPG, a novel visual Autoregressive model that enables Randomized Parallel Generation, addressing the inherent limitations of conventional raster-order approaches, whi…
PISA: Piecewise Sparse Attention Is Wiser for Efficient Diffusion Transformers
Haopeng Li, Shitong Shao, Wenliang Zhong +4
Diffusion Transformers are fundamental for video and image generation, but their efficiency is bottlenecked by the quadratic complexity of attention. While block sparse attention a…
Can Large Language Models Analyze Graphs like Professionals? A Benchmark, Datasets and Models
Xin Li, Weize Chen, Qizhi Chu +9
The need to analyze graphs is ubiquitous across various fields, from social networks to biological research and recommendation systems. Therefore, enabling the ability of large lan…