activity
20242026
collaborators

6 papers

cs.LG2026

Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training

Woojeong Kim, Ziyi Yang, Jing Nathan Yan +1

Reinforcement learning (RL) is the dominant paradigm for post-training large language models. However, in the online, on-policy setting, rollout generation dominates the computatio…

cs.CL2025

LiPO: Listwise Preference Optimization through Learning-to-Rank

Tianqi Liu, Zhen Qin, Junru Wu +9

Aligning language models (LMs) with curated human feedback is critical to control their behaviors in real-world applications. Several recent policy optimization methods, such as DP…

cs.CV2024

Video Summarization: Towards Entity-Aware Captions

Hammad A. Ayyubi, Tianqi Liu, Arsha Nagrani +7

Existing popular video captioning benchmarks and models deal with generic captions devoid of specific person, place or organization named entities. In contrast, news videos present…

cs.CL2024

Multilingual Fine-Grained News Headline Hallucination Detection

Jiaming Shen, Tianqi Liu, Jialu Liu +4

The popularity of automated news headline generation has surged with advancements in pre-trained language models. However, these models often suffer from the ``hallucination'' prob…

cs.CL2024

Predicting Text Preference Via Structured Comparative Reasoning

Jing Nathan Yan, Tianqi Liu, Justin T Chiu +9

Comparative reasoning plays a crucial role in text preference prediction; however, large language models (LLMs) often demonstrate inconsistencies in their reasoning. While approach…

cs.CL2024

PLaD: Preference-based Large Language Model Distillation with Pseudo-Preference Pairs

Rongzhi Zhang, Jiaming Shen, Tianqi Liu +7

Large Language Models (LLMs) have exhibited impressive capabilities in various tasks, yet their vast parameter sizes restrict their applicability in resource-constrained settings.…