14 papers
BiasTrace: Linking Reasoning Behaviours to Biased Outputs in LLMs
Varsha Ramineni, Hossein A. Rahmani, Jerome Ramos +2
LLMs exhibit social biases that can produce inaccurate and discriminatory inferences, posing risks in high-stakes applications. While prior work has made progress in measuring and…
InnoEval: On Research Idea Evaluation as a Knowledge-Grounded, Multi-Perspective Reasoning Problem
Shuofei Qiao, Yunxiang Wei, Xuehai Wang +10
The rapid evolution of Large Language Models has catalyzed a surge in scientific idea production, yet this leap has not been accompanied by a matching advance in idea evaluation. T…
Automated Rubrics for Reliable Evaluation of Medical Dialogue Systems
Yinzhu Chen, Abdine Maiga, Hossein A. Rahmani +1
Large Language Models (LLMs) are increasingly used for clinical decision support, where hallucinations and unsafe suggestions may pose direct risks to patient safety. These risks a…
Interplay: Training Independent Simulators for Reference-Free Conversational Recommendation
Jerome Ramos, Feng Xia, Xi Wang +4
Training conversational recommender systems (CRS) requires extensive dialogue data, which is challenging to collect at scale. To address this, researchers have used simulated user-…
Beyond Output Critique: Self-Correction via Task Distillation
Hossein A. Rahmani, Mengting Wan, Pei Zhou +4
Large language models (LLMs) have shown promising self-correction abilities, where iterative refinement improves the quality of generated responses. However, most existing approach…
Self-Correcting Large Language Models: Generation vs. Multiple Choice
Hossein A. Rahmani, Satyapriya Krishna, Xi Wang +2
Large language models have recently demonstrated remarkable abilities to self-correct their responses through iterative refinement, often referred to as self-consistency or self-re…