4 papers
VeriSim: A Configurable Framework for Evaluating Medical AI Under Realistic Patient Noise
Sina Mansouri, Mohit Marvania, Vibhavari Ashok Shihorkar +5
Medical large language models (LLMs) achieve impressive performance on standardized benchmarks, yet these evaluations fail to capture the complexity of real clinical encounters whe…
WISE-Flow: Workflow-Induced Structured Experience for Self-Evolving Conversational Service Agents
Yuqing Zhou, Zhuoer Wang, Jie Yuan +4
Large language model (LLM)-based agents are widely deployed in user-facing services but remain error-prone in new tasks, tend to repeat the same failure patterns, and show substant…
Fighting Spurious Correlations in Text Classification via a Causal Learning Perspective
Yuqing Zhou, Ziwei Zhu
In text classification tasks, models often rely on spurious correlations for predictions, incorrectly associating irrelevant features with the target labels. This issue limits the…
Navigating the Shortcut Maze: A Comprehensive Analysis of Shortcut Learning in Text Classification by Language Models
Yuqing Zhou, Ruixiang Tang, Ziyu Yao +1
Language models (LMs), despite their advances, often depend on spurious correlations, undermining their accuracy and generalizability. This study addresses the overlooked impact of…