11 papers
Forecast for the detectability of patchy hydrogen reionization in WEAVE-QSO measurements of the Lyman- forest power spectrum at redshift
Ke Ma, James S. Bolton, Vid Iršič +9
We present the first detailed forecasts for the detectability of patchy hydrogen reionization in the one-dimensional Ly forest power spectrum to be measured by the WEAVE-QSO sur…
Compress the Easy, Explore the Hard: Difficulty-Aware Entropy Regularization for Efficient LLM Reasoning
Qin-Wen Luo, Sheng Ren, Xiang Chen +4
Chain-of-Thought (CoT) has substantially empowered Large Language Models (LLMs) to tackle complex reasoning tasks, yet the verbose nature of explicit reasoning steps incurs prohibi…
Who Deserves the Reward? SHARP: Shapley Credit-based Optimization for Multi-Agent System
Yanming Li, Xuelin Zhang, WenJie Lu +11
Integrating Large Language Models (LLMs) with external tools via multi-agent systems offers a promising new paradigm for decomposing and solving complex problems. However, training…
Mitigating Safety Tax via Distribution-Grounded Refinement in Large Reasoning Models
Yingsha Xie, Tiansheng Huang, Enneng Yang +5
Safety alignment incurs safety tax that perturbs a large reasoning model's (LRM) general reasoning ability. Existing datasets used for safety alignment for an LRM are usually const…
SafeThinker: Reasoning about Risk to Deepen Safety Beyond Shallow Alignment
Xianya Fang, Xianying Luo, Yadong Wang +8
Despite the intrinsic risk-awareness of Large Language Models (LLMs), current defenses often result in shallow safety alignment, rendering models vulnerable to disguised attacks (e…
Panacea: Mitigating Harmful Fine-tuning for Large Language Models via Post-fine-tuning Perturbation
Yibo Wang, Tiansheng Huang, Li Shen +6
Harmful fine-tuning attack introduces significant security risks to the fine-tuning services. Main-stream defenses aim to vaccinate the model such that the later harmful fine-tunin…