collaborators

12 papers

cs.LG2026

Policy Improvement Reinforcement Learning

Huaiyang Wang, Xiaojie Li, Xiaohan Wang +10

Reinforcement learning has become a central post-training paradigm for improving LLM and agent capabilities. Yet existing RL post-training methods share a common blind spot: they c…

cs.LG2025

FRoD: Full-Rank Efficient Fine-Tuning with Rotational Degrees for Fast Convergence

Guoan Wan, Tianyu Chen, Fangzheng Feng +2

Parameter-efficient fine-tuning (PEFT) methods have emerged as a practical solution for adapting large foundation models to downstream tasks, reducing computational and memory cost…

cs.CV2025

Towards Long-window Anchoring in Vision-Language Model Distillation

Haoyi Zhou, Shuo Li, Tianyu Chen +3

While large vision-language models (VLMs) demonstrate strong long-context understanding, their prevalent small branches fail on linguistics-photography alignment for a limited wind…

cs.CL2025

Fine-Tuned LLMs Know They Don't Know: A Parameter-Efficient Approach to Recovering Honesty

Zeyu Shi, Ziming Wang, Tianyu Chen +4

The honesty of Large Language Models (LLMs) is increasingly important for safe deployment in high-stakes domains. However, this crucial trait is severely undermined by supervised f…

cs.AI2025

Accelerate Scaling of LLM Finetuning via Quantifying the Coverage and Depth of Instruction Set

Chengwei Wu, Li Du, Hanyu Zhao +4

Scaling the amount of data used for supervied fine-tuning(SFT) does not guarantee the proportional gains in model performance, highlighting a critical need to understand what makes…

cs.LG2025

LLM-Mixer: Multiscale Mixing in LLMs for Time Series Forecasting

Md Kowsher, Md. Shohanur Islam Sobuj, Nusrat Jahan Prottasha +3

Time series forecasting remains a challenging task, particularly in the context of complex multiscale temporal patterns. This study presents LLM-Mixer, a framework that improves fo…