collaborators

7 papers

cs.CL2026

TIAR: Trajectory-Informed Advantage Reweighting for LLM Abstention Learning

Muyu Pan, Shu Zhao, Nan Zhang +4

This paper investigates large language model (LLM) abstention learning, specifically using ternary reward, which incentivize truthfulness in large language models. This paper exten…

cs.CL2026

From Correctness to Preference: A Framework for Personalized Agentic Reinforcement Learning

Ranxu zhang, zeyang li, Jiacheng Huang +5

Agentic reinforcement learning (Agentic RL) has achieved strong progress in tasks with clear success signals. However, many real-world agent applications require user-conditioned b…

cs.LG2026

Why Do Reasoning Models Lose Coverage? The Role of Data and Forks in the Road

Ngoc-Hieu Nguyen, Parshin Shojaee, Phuc Minh Nguyen +4

Recent progress in large language models has led to the emergence of reasoning models, which have shown strong performance on complex tasks through specialized fine-tuning procedur…

cs.LG2026

When Reasoning Meets Compression: Understanding the Effects of LLMs Compression on Large Reasoning Models

Nan Zhang, Eugene Kwek, Yusen Zhang +3

Compression methods, including quantization, distillation, and pruning, improve the computational efficiency of large reasoning models (LRMs). However, existing studies either fail…

cs.LG2026

QuantLRM: Quantization of Large Reasoning Models via Fine-Tuning Signals

Nan Zhang, Eugene Kwek, Yusen Zhang +4

Weight-only quantization is important for compressing Large Language Models (LLMs). Inspired by the spirit of classical magnitude pruning, we study whether the magnitude of weight…

cs.CL2025

TaxoAdapt: Aligning LLM-Based Multidimensional Taxonomy Construction to Evolving Research Corpora

Priyanka Kargupta, Nan Zhang, Yunyi Zhang +3

The rapid evolution of scientific fields introduces challenges in organizing and retrieving scientific literature. While expert-curated taxonomies have traditionally addressed this…