collaborators

15 papers

cs.LG2026

Inverse RL Helps Align AI by Imitating Humans

Michał Wiliński, Liu Leqi, Chirag Nagpal

Language model alignment aims to make model behavior reliably reflect desirable properties such as helpfulness, safety, and instruction following. Current approaches typically use…

cs.CL2026

The Severance Problem: LLMs are Unaware of the Person Beyond the Prompt

Dor Litvak, Liu Leqi

The paper identifies the "Severance Problem"—the lack of an explicit representation of a user beyond the prompt—in current LLM-based personal assistants, and proposes a "Severance…

cs.AI2026

Evaluating Stochasticity in Deep Research Agents

Haotian Zhai, Elias Stengel-Eskin, Pratik Patil +1

Deep Research Agents (DRAs) are promising agentic systems that gather and synthesize information to support research across domains such as financial decision-making, medical analy…

cs.LG2026

Learning Robust Reasoning through Guided Adversarial Self-Play

Shuozhe Li, Vaishnav Tadiparthi, Kwonjoon Lee +6

Reinforcement learning from verifiable rewards (RLVR) produces strong reasoning models, yet they can fail catastrophically when the conditioning context is fallible (e.g., corrupte…

cs.LG2026

ExPO: Unlocking Hard Reasoning with Self-Explanation-Guided Reinforcement Learning

Ruiyang Zhou, Shuozhe Li, Amy Zhang +1

Self-improvement via RL often fails on complex reasoning tasks because GRPO-style post-training methods rely on the model's initial ability to generate positive samples. Without gu…

cs.CL2025

Position: Thematic Analysis of Unstructured Clinical Transcripts with Large Language Models

Seungjun Yi, Joakim Nguyen, Terence Lim +8

This position paper examines how large language models (LLMs) can support thematic analysis of unstructured clinical transcripts, a widely used but resource-intensive method for un…