5 papers
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs
Di He, Songjun Tu, Keyu Wang +2
Learning rate configuration is a fundamental aspect of modern deep learning. The prevailing practice of applying a uniform learning rate across all layers overlooks the structural…
Demystifying the Roles of LLM Layers in Retrieval, Knowledge, and Reasoning
Xinyuan Song, Keyu Wang, PengXiang Li +2
Recent studies suggest that the deeper layers of Large Language Models (LLMs) contribute little to representation learning and can often be removed without significant performance…
CORE: Code-based Inverse Self-Training Framework with Graph Expansion for Virtual Agents
Keyu Wang, Bingchen Miao, Wendong Bu +7
The development of Multimodal Virtual Agents has made significant progress through the integration of Multimodal Large Language Models. However, mainstream training paradigms face…
When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs
Keyu Wang, Tian Lyu, Guinan Su +4
Layer pruning has emerged as a widely adopted technique for improving the efficiency of large language models (LLMs). Although existing methods demonstrate strong performance reten…
Mitigating Downstream Model Risks via Model Provenance
Keyu Wang, Abdullah Norozi Iranzad, Scott Schaffter +3
Research and industry are rapidly advancing the innovation and adoption of foundation model-based systems, yet the tools for managing these models have not kept pace. Understanding…