12 papers
Policy Improvement Reinforcement Learning
Huaiyang Wang, Xiaojie Li, Xiaohan Wang +10
Reinforcement learning has become a central post-training paradigm for improving LLM and agent capabilities. Yet existing RL post-training methods share a common blind spot: they c…
FRoD: Full-Rank Efficient Fine-Tuning with Rotational Degrees for Fast Convergence
Guoan Wan, Tianyu Chen, Fangzheng Feng +2
Parameter-efficient fine-tuning (PEFT) methods have emerged as a practical solution for adapting large foundation models to downstream tasks, reducing computational and memory cost…
Towards Long-window Anchoring in Vision-Language Model Distillation
Haoyi Zhou, Shuo Li, Tianyu Chen +3
While large vision-language models (VLMs) demonstrate strong long-context understanding, their prevalent small branches fail on linguistics-photography alignment for a limited wind…
Fine-Tuned LLMs Know They Don't Know: A Parameter-Efficient Approach to Recovering Honesty
Zeyu Shi, Ziming Wang, Tianyu Chen +4
The honesty of Large Language Models (LLMs) is increasingly important for safe deployment in high-stakes domains. However, this crucial trait is severely undermined by supervised f…
Accelerate Scaling of LLM Finetuning via Quantifying the Coverage and Depth of Instruction Set
Chengwei Wu, Li Du, Hanyu Zhao +4
Scaling the amount of data used for supervied fine-tuning(SFT) does not guarantee the proportional gains in model performance, highlighting a critical need to understand what makes…
LLM-Mixer: Multiscale Mixing in LLMs for Time Series Forecasting
Md Kowsher, Md. Shohanur Islam Sobuj, Nusrat Jahan Prottasha +3
Time series forecasting remains a challenging task, particularly in the context of complex multiscale temporal patterns. This study presents LLM-Mixer, a framework that improves fo…