activity
20242026
collaborators

8 papers

cs.CV2026

HSMLA: Hierarchical Softmax Multi-scale Linear Attention for Efficient Vision Transformers

Dong Liu, Yanxuan Yu, Renata Borovica-Gajic +2

Vision transformers face significant computational overheads in high-resolution dense prediction due to the quadratic complexity of self-attention. Linear attention offers efficien…

cs.CL2025

Probabilistic Token Alignment for Large Language Model Fusion

Runjia Zeng, James Chenhao Liang, Cheng Han +8

Training large language models (LLMs) from scratch can yield models with unique functionalities and strengths, but it is costly and often leads to redundant capabilities. A more co…

cs.LG2025

MEPT: Mixture of Expert Prompt Tuning as a Manifold Mapper

Runjia Zeng, Guangyan Sun, Qifan Wang +8

Considering deep neural networks as manifold mappers, the pretrain-then-fine-tune paradigm can be interpreted as a two-stage process: pretrain establishes a broad knowledge base, a…

cs.CV2024

Visual Fourier Prompt Tuning

Runjia Zeng, Cheng Han, Qifan Wang +5

With the scale of vision Transformer-based models continuing to grow, finetuning these large-scale pretrained models for new tasks has become increasingly parameter-intensive. Visu…

cs.LG2024

Inertial Confinement Fusion Forecasting via Large Language Models

Mingkai Chen, Taowen Wang, Shihui Cao +10

Controlled fusion energy is deemed pivotal for the advancement of human civilization. In this study, we introduce , a novel integration of Large Language Models (…

physics.comp-ph2024

Diff-PIC: Revolutionizing Particle-In-Cell Nuclear Fusion Simulation with Diffusion Models

Chuan Liu, Chunshu Wu, Shihui Cao +8

The rapid development of AI highlights the pressing need for sustainable energy, a critical global challenge for decades. Nuclear fusion, generally seen as an ultimate solution, ha…