9 papers
Prediction Under Imperfect Compression: A Theory of Approximate MDL
Qian Li, Xinyu Mao, Shang-Hua Teng +1
Minimum Description Length (MDL) formalizes the principle of Occam's razor by optimizing the total description length: . Fo…
ChunkFT: Byte-Streamed Optimization for Memory-Efficient Full Fine-Tuning
Yongkang Liu, Zijing Wang, Mengjie Zhao +7
This work presents \textsc{ChunkFT}, a memory-efficient fine-tuning framework that reformulates full-parameter fine-tuning around a dynamically activated working set. \textsc{Chunk…
SMoA: Spectrum Modulation Adapter for Parameter-Efficient Fine-Tuning
Yongkang Liu, Xing Li, Mengjie Zhao +7
As the number of model parameters increases, parameter-efficient fine-tuning (PEFT) has become the go-to choice for tailoring pre-trained large language models. Low-rank Adaptation…
expo: Exploration-prioritized policy optimization via adaptive kl regulation and gaussian curriculum sampling
Mingxiong Lin, Zhangquan Gong, Maowen Tang +6
Reinforcement Learning with Verifiable Rewards (RLVR) has become the standard paradigm for LLM mathematical reasoning, where Group Relative Policy Optimization (GRPO) serves as the…
fg-expo: Frontier-guided exploration-prioritized policy optimization via adaptive kl and gaussian curriculum
Mingxiong Lin, Zhangquan Gong, Maowen Tang +6
Reinforcement Learning with Verifiable Rewards (RLVR) has become the standard paradigm for LLM mathematical reasoning, with Group Relative Policy Optimization (GRPO) serving as the…
Reducing Detail Hallucinations in Long-Context Regulatory Understanding via Targeted Preference Optimization
Yang Liu, Bin Chong, Yuhan Lin +7
Large language models (LLMs) frequently produce \emph{detail hallucinations} when processing long regulatory documents, including subtle errors in threshold values, units, scopes,…