collaborators

6 papers

cs.LG2026

Bottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal Policies

Yuqiao Tan, Minzheng Wang, Shizhu He +6

Existing reinforcement learning (RL) approaches treat large language models (LLMs) as a unified policy, overlooking their internal mechanisms. In this paper, we decompose the LLM-b…

cs.AI2026

Data Science and Technology Towards AGI Part I: Tiered Data Management

Yudong Wang, Zixuan Fu, Hengyu Zhao +14

The development of artificial intelligence can be viewed as an evolution of data-driven learning paradigms, with successive shifts in data organization and utilization continuously…

cs.CL2026

Factuality on Demand: Controlling the Factuality-Informativeness Trade-off in Text Generation

Ziwei Gong, Yanda Chen, Julia Hirschberg +4

Large language models (LLMs) encode knowledge with varying degrees of confidence. When responding to queries, models face an inherent trade-off: they can generate responses that ar…

cs.CL2026

Small Updates, Big Doubts: Does Parameter-Efficient Fine-tuning Enhance Hallucination Detection ?

Xu Hu, Yifan Zhang, Songtao Wei +4

Parameter-efficient fine-tuning (PEFT) methods are widely used to adapt large language models (LLMs) to downstream tasks and are often assumed to improve factual correctness. Howev…

cs.CL2025

An Analysis of Large Language Models for Simulating User Responses in Surveys

Ziyun Yu, Yiru Zhou, Chen Zhao +1

Using Large Language Models (LLMs) to simulate user opinions has received growing attention. Yet LLMs, especially trained with reinforcement learning from human feedback (RLHF), ar…

cs.CL2025

Leaps Beyond the Seen: Reinforced Reasoning Augmented Generation for Clinical Notes

Lo Pang-Yun Ting, Chengshuai Zhao, Yu-Hua Zeng +3

Clinical note generation aims to produce free-text summaries of a patient's condition and diagnostic process, with discharge instructions being a representative long-form example.…