activity
20242026
collaborators

14 papers

cs.AI2026

BiasTrace: Linking Reasoning Behaviours to Biased Outputs in LLMs

Varsha Ramineni, Hossein A. Rahmani, Jerome Ramos +2

LLMs exhibit social biases that can produce inaccurate and discriminatory inferences, posing risks in high-stakes applications. While prior work has made progress in measuring and…

cs.CL2026

InnoEval: On Research Idea Evaluation as a Knowledge-Grounded, Multi-Perspective Reasoning Problem

Shuofei Qiao, Yunxiang Wei, Xuehai Wang +10

The rapid evolution of Large Language Models has catalyzed a surge in scientific idea production, yet this leap has not been accompanied by a matching advance in idea evaluation. T…

cs.CL2026

Automated Rubrics for Reliable Evaluation of Medical Dialogue Systems

Yinzhu Chen, Abdine Maiga, Hossein A. Rahmani +1

Large Language Models (LLMs) are increasingly used for clinical decision support, where hallucinations and unsafe suggestions may pose direct risks to patient safety. These risks a…

cs.AI2026

Interplay: Training Independent Simulators for Reference-Free Conversational Recommendation

Jerome Ramos, Feng Xia, Xi Wang +4

Training conversational recommender systems (CRS) requires extensive dialogue data, which is challenging to collect at scale. To address this, researchers have used simulated user-…

cs.AI2026

Beyond Output Critique: Self-Correction via Task Distillation

Hossein A. Rahmani, Mengting Wan, Pei Zhou +4

Large language models (LLMs) have shown promising self-correction abilities, where iterative refinement improves the quality of generated responses. However, most existing approach…

cs.CL2025

Self-Correcting Large Language Models: Generation vs. Multiple Choice

Hossein A. Rahmani, Satyapriya Krishna, Xi Wang +2

Large language models have recently demonstrated remarkable abilities to self-correct their responses through iterative refinement, often referred to as self-consistency or self-re…