5 papers
MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokens
Yu Chen, Runkai Chen, Sheng Yi +9
Long-term memory is a cornerstone of human intelligence. Enabling AI to process lifetime-scale information remains a long-standing pursuit in the field. Due to the constraints of f…
TempoFit: Plug-and-Play Layer-Wise Temporal KV Memory for Long-Horizon Vision-Language-Action Manipulation
Jun Sun, Boyu Yang, Jiahao Zhang +7
Pretrained Vision-Language-Action (VLA) policies have achieved strong single-step manipulation, but their inference remains largely memoryless, which is brittle in non-Markovian lo…
Backdoor4Good: Benchmarking Beneficial Uses of Backdoors in LLMs
Yige Li, Wei Zhao, Zhe Li +6
Backdoor mechanisms have traditionally been studied as security threats that compromise the integrity of machine learning models. However, the same mechanism -- the conditional act…
Induced Numerical Instability: Hidden Costs in Multimodal Large Language Models
Wai Tuck Wong, Jun Sun, Arunesh Sinha
The use of multimodal large language models has become widespread, and as such the study of these models and their failure points has become of utmost importance. We study a novel…
Following the Teacher's Footsteps: Scheduled Checkpoint Distillation for Domain-Specific LLMs
Cheng Feng, Chaoliang Zhong, Jun Sun +1
Large language models (LLMs) are challenging to deploy for domain-specific tasks due to their massive scale. While distilling a fine-tuned LLM into a smaller student model is a pro…