2 papers
cs.LG2026
NNiT: Width-Agnostic Neural Network Generation with Structurally Aligned Weight Spaces
Jiwoo Kim, Swarajh Mehta, Hao-Lun Hsu +3
Generative modeling of neural network parameters is often tied to architectures because standard parameter representations rely on known weight-matrix dimensions. Generation is fur…
cs.DC2025
Practical offloading for fine-tuning LLM on commodity GPU via learned sparse projectors
Siyuan Chen, Zhuofeng Wang, Zelong Guan +2
Fine-tuning large language models (LLMs) requires significant memory, often exceeding the capacity of a single GPU. A common solution to this memory challenge is offloading compute…