bidirectional generation 1hybrid attention 1language model fine-tuning 1masked diffusion language modeling 1pretrained autoregressive models 1
From the 1 of 21 linked papers with an AI index.
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
ENCORE: Entropy-guided Reward Composition for Multi-head Safety Reward Models
Xiaomin Li, Xupeng Chen, Jingxuan Fan +2
The safety alignment of large language models (LLMs) often relies on reinforcement learning from human feedback (RLHF), which requires human annotations to construct preference dat…
cs.CL2025
When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs
Xiaomin Li, Zhou Yu, Zhiwei Zhang +5
Reasoning-enhanced large language models (RLLMs), whether explicitly trained for reasoning or prompted via chain-of-thought (CoT), have achieved state-of-the-art performance on man…
cs.CL2025
Instruction Tuning and CoT Prompting for Contextual Medical QA with LLMs
Chenqian Le, Ziheng Gong, Chihang Wang +3
Large language models (LLMs) have shown great potential in medical question answering (MedQA), yet adapting them to biomedical reasoning remains challenging due to domain-specific…