4 papers
EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture
Bowen Duan, Cong Guo, Chiyue Wei +8
Large Language Models (LLMs) have achieved impressive performance across diverse domains but remain inefficient during the autoregressive decoding phase. Unlike the prefill stage,…
How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors
Yifan Xu, Junren Chen, Yifan Chen
Reinforcement learning with verifiable rewards (RLVR) recently thrives in large language model (LLM) reasoning tasks. However, the reward sparsity and the long reasoning horizon ma…
When Human Preferences Flip: An Instance-Dependent Robust Loss for RLHF
Yifan Xu, Xichen Ye, Yifan Chen +1
Quality of datasets plays an important role in large language model (LLM) alignment. In collecting human feedback, however, preference flipping is ubiquitous and causes corruption…
Optimizing LVLMs with On-Policy Data for Effective Hallucination Mitigation
Chengzhi Yu, Yifan Xu, Yifan Chen +1
Recently, large vision-language models (LVLMs) have risen to be a promising approach for multimodal tasks. However, principled hallucination mitigation remains a critical challenge…