collaborators

5 papers

stat.ML2026

When Does Trajectory-Level Supervision Permit Efficient Offline Reinforcement Learning?

Xuanfei Ren, Tengyang Xie

Offline reinforcement learning is typically analyzed under process-level reward supervision, yet many sequential decision datasets record only trajectory-level outcomes. We develop…

cs.LG2026

POLCA: Stochastic Generative Optimization with LLM

Xuanfei Ren, Allen Nie, Tengyang Xie +1

Optimizing complex systems, ranging from LLM prompts to multi-turn agents, traditionally requires labor-intensive manual iteration. We formalize this challenge as a stochastic gene…

cs.CL2025

I Learn Better If You Speak My Language: Understanding the Superior Performance of Fine-Tuning Large Language Models with LLM-Generated Responses

Xuan Ren, Biao Wu, Lingqiao Liu

This paper explores an intriguing observation: fine-tuning a large language model (LLM) with responses generated by a LLM often yields better results than using responses generated…

cs.CL2025

Efficient Response Generation Strategy Selection for Fine-Tuning Large Language Models Through Self-Aligned Perplexity

Xuan Ren, Qi Chen, Lingqiao Liu

Fine-tuning large language models (LLMs) typically relies on producing large sets of input-output pairs. Yet for a given question, there can be many valid outputs. In practice, the…

cs.CL2025

You Can Generate It Again: Data-to-Text Generation with Verification and Correction Prompting

Xuan Ren, Zeyu Zhang, Lingqiao Liu

Small language models like T5 excel in generating high-quality text for data-to-text tasks, offering adaptability and cost-efficiency compared to Large Language Models (LLMs). Howe…