activity
20242026
collaborators

9 papers

cs.LG2026

Prediction Under Imperfect Compression: A Theory of Approximate MDL

Qian Li, Xinyu Mao, Shang-Hua Teng +1

Minimum Description Length (MDL) formalizes the principle of Occam's razor by optimizing the total description length: . Fo…

cs.LG2026

ChunkFT: Byte-Streamed Optimization for Memory-Efficient Full Fine-Tuning

Yongkang Liu, Zijing Wang, Mengjie Zhao +7

This work presents \textsc{ChunkFT}, a memory-efficient fine-tuning framework that reformulates full-parameter fine-tuning around a dynamically activated working set. \textsc{Chunk…

cs.LG2026

SMoA: Spectrum Modulation Adapter for Parameter-Efficient Fine-Tuning

Yongkang Liu, Xing Li, Mengjie Zhao +7

As the number of model parameters increases, parameter-efficient fine-tuning (PEFT) has become the go-to choice for tailoring pre-trained large language models. Low-rank Adaptation…

cs.AI2026

expo: Exploration-prioritized policy optimization via adaptive kl regulation and gaussian curriculum sampling

Mingxiong Lin, Zhangquan Gong, Maowen Tang +6

Reinforcement Learning with Verifiable Rewards (RLVR) has become the standard paradigm for LLM mathematical reasoning, where Group Relative Policy Optimization (GRPO) serves as the…

cs.LG2026

fg-expo: Frontier-guided exploration-prioritized policy optimization via adaptive kl and gaussian curriculum

Mingxiong Lin, Zhangquan Gong, Maowen Tang +6

Reinforcement Learning with Verifiable Rewards (RLVR) has become the standard paradigm for LLM mathematical reasoning, with Group Relative Policy Optimization (GRPO) serving as the…

cs.SI2026

Reducing Detail Hallucinations in Long-Context Regulatory Understanding via Targeted Preference Optimization

Yang Liu, Bin Chong, Yuhan Lin +7

Large language models (LLMs) frequently produce \emph{detail hallucinations} when processing long regulatory documents, including subtle errors in threshold values, units, scopes,…