2 citations · 3 across the 7 of their papers we have counts for
7 papers
Spurious Advantage Hidden in GRPO
Jiamian Wang, Samyadeep Basu, Koustava Goswami +2
Group Relative Policy Optimization (GRPO) is widely studied for reinforcement learning with verifiable rewards, where its advantage estimator assigns each rollout a magnitude from…
HindSearch: Trajectory-Level Hindsight Critique for Search-Augmented Reinforcement Learning
Haowei Liu, Jiamian Wang, Hsin-Tai Wu +2
Search-augmented LM agents are typically trained with a binary exact-match reward, which throws away most of what a failed trajectory tells us about why it failed. We introduce Hin…
Does Reasoning Introduce Bias? A Study of Social Bias Evaluation and Mitigation in LLM Reasoning
Xuyang Wu, Jinming Nian, Ting-Ruen Wei +3
Recent advances in large language models (LLMs) have enabled automatic generation of chain-of-thought (CoT) reasoning, leading to strong performance on tasks such as math and code.…
RaCT: Ranking-aware Chain-of-Thought Optimization for LLMs
Haowei Liu, Xuyang Wu, Guohao Sun +2
In information retrieval, large language models (LLMs) have demonstrated remarkable potential in text reranking tasks by leveraging their sophisticated natural language understandi…
Does RAG Introduce Unfairness in LLMs? Evaluating Fairness in Retrieval-Augmented Generation Systems
Xuyang Wu, Shuowei Li, Hsin-Tai Wu +2
Retrieval-Augmented Generation (RAG) has recently gained significant attention for its enhanced ability to integrate external knowledge sources into open-domain question answering…
Evaluating Fairness in Large Vision-Language Models Across Diverse Demographic Attributes and Prompts
Xuyang Wu, Yuan Wang, Hsin-Tai Wu +2
Large vision-language models (LVLMs) have recently achieved significant progress, demonstrating strong capabilities in open-world visual understanding. However, it is not yet clear…