6 papers
Bottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal Policies
Yuqiao Tan, Minzheng Wang, Shizhu He +6
Existing reinforcement learning (RL) approaches treat large language models (LLMs) as a unified policy, overlooking their internal mechanisms. In this paper, we decompose the LLM-b…
Data Science and Technology Towards AGI Part I: Tiered Data Management
Yudong Wang, Zixuan Fu, Hengyu Zhao +14
The development of artificial intelligence can be viewed as an evolution of data-driven learning paradigms, with successive shifts in data organization and utilization continuously…
Factuality on Demand: Controlling the Factuality-Informativeness Trade-off in Text Generation
Ziwei Gong, Yanda Chen, Julia Hirschberg +4
Large language models (LLMs) encode knowledge with varying degrees of confidence. When responding to queries, models face an inherent trade-off: they can generate responses that ar…
Small Updates, Big Doubts: Does Parameter-Efficient Fine-tuning Enhance Hallucination Detection ?
Xu Hu, Yifan Zhang, Songtao Wei +4
Parameter-efficient fine-tuning (PEFT) methods are widely used to adapt large language models (LLMs) to downstream tasks and are often assumed to improve factual correctness. Howev…
An Analysis of Large Language Models for Simulating User Responses in Surveys
Ziyun Yu, Yiru Zhou, Chen Zhao +1
Using Large Language Models (LLMs) to simulate user opinions has received growing attention. Yet LLMs, especially trained with reinforcement learning from human feedback (RLHF), ar…
Leaps Beyond the Seen: Reinforced Reasoning Augmented Generation for Clinical Notes
Lo Pang-Yun Ting, Chengshuai Zhao, Yu-Hua Zeng +3
Clinical note generation aims to produce free-text summaries of a patient's condition and diagnostic process, with discharge instructions being a representative long-form example.…