8 papers
HSMLA: Hierarchical Softmax Multi-scale Linear Attention for Efficient Vision Transformers
Dong Liu, Yanxuan Yu, Renata Borovica-Gajic +2
Vision transformers face significant computational overheads in high-resolution dense prediction due to the quadratic complexity of self-attention. Linear attention offers efficien…
Probabilistic Token Alignment for Large Language Model Fusion
Runjia Zeng, James Chenhao Liang, Cheng Han +8
Training large language models (LLMs) from scratch can yield models with unique functionalities and strengths, but it is costly and often leads to redundant capabilities. A more co…
MEPT: Mixture of Expert Prompt Tuning as a Manifold Mapper
Runjia Zeng, Guangyan Sun, Qifan Wang +8
Considering deep neural networks as manifold mappers, the pretrain-then-fine-tune paradigm can be interpreted as a two-stage process: pretrain establishes a broad knowledge base, a…
Visual Fourier Prompt Tuning
Runjia Zeng, Cheng Han, Qifan Wang +5
With the scale of vision Transformer-based models continuing to grow, finetuning these large-scale pretrained models for new tasks has become increasingly parameter-intensive. Visu…
Inertial Confinement Fusion Forecasting via Large Language Models
Mingkai Chen, Taowen Wang, Shihui Cao +10
Controlled fusion energy is deemed pivotal for the advancement of human civilization. In this study, we introduce , a novel integration of Large Language Models (…
Diff-PIC: Revolutionizing Particle-In-Cell Nuclear Fusion Simulation with Diffusion Models
Chuan Liu, Chunshu Wu, Shihui Cao +8
The rapid development of AI highlights the pressing need for sustainable energy, a critical global challenge for decades. Nuclear fusion, generally seen as an ultimate solution, ha…