7 papers
MMR-GRPO: Accelerating GRPO-Style Training through Diversity-Aware Reward Reweighting
Kangda Wei, Ruihong Huang
Group Relative Policy Optimization (GRPO) has become a standard approach for training mathematical reasoning models; however, its reliance on multiple completions per prompt makes…
Sycophantic Praise: Evaluating Excessive Praise in Language Models
Daniel Vennemeyer, Phan Anh Duong, Meryl Ye +2
Sycophancy in language models is typically studied as excessive agreement or validation, while explicit praise and flattery have received comparatively little attention. We argue t…
Mitigating Gender Bias via Fostering Exploratory Thinking in LLMs
Kangda Wei, Hasnat Md Abdullah, Ruihong Huang
Large Language Models (LLMs) often exhibit gender bias, resulting in unequal treatment of male and female subjects across different contexts. To address this issue, we propose a no…
CliME: Evaluating Multimodal Climate Discourse on Social Media and the Climate Alignment Quotient (CAQ)
Abhilekh Borah, Hasnat Md Abdullah, Kangda Wei +1
The rise of Large Language Models (LLMs) has raised questions about their ability to understand climate-related contexts. Though climate change dominates social media, analyzing it…
LegalCore: A Dataset for Event Coreference Resolution in Legal Documents
Kangda Wei, Xi Shi, Jonathan Tong +5
Recognizing events and their coreferential mentions in a document is essential for understanding semantic meanings of text. The existing research on event coreference resolution is…
PreMind: Multi-Agent Video Understanding for Advanced Indexing of Presentation-style Videos
Kangda Wei, Zhengyu Zhou, Bingqing Wang +4
In recent years, online lecture videos have become an increasingly popular resource for acquiring new knowledge. Systems capable of effectively understanding/indexing lecture video…