Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning
Guochao Jiang, Jingyi Song, Guofeng Quan +3
Reinforcement Learning has become a standard paradigm for aligning Large Language Models with human intent and task requirements. While Group Relative Policy Optimization offers an…
cs.CL2025
RASD: Retrieval-Augmented Speculative Decoding
Guofeng Quan, Wenfeng Feng, Chuzhan Hao +3
Speculative decoding accelerates inference in large language models (LLMs) by generating draft tokens for target model verification. Current approaches for obtaining draft tokens r…
cs.CL2024
Mixture-of-LoRAs: An Efficient Multitask Tuning for Large Language Models
Wenfeng Feng, Chuzhan Hao, Yuewei Zhang +2
Instruction Tuning has the potential to stimulate or enhance specific capabilities of large language models (LLMs). However, achieving the right balance of data is crucial to preve…