5 papers · 1 filter
TimeSage-MT: A Multi-Turn Benchmark for Evaluating Agentic Time Series Reasoning
Yaxuan Kong, Qingren Yao, Yuqi Nie +7
Time series data inform critical decisions across many real-world domains. While large language model (LLM) agents can analyze data through natural language and tools, it remains u…
CoCoReviewBench: A Completeness- and Correctness-Oriented Benchmark for AI Reviewers
Hexuan Deng, Xiaopeng Ke, Yichen Li +6
Despite the rapid development of AI reviewers, evaluating such systems remains challenging: metrics favor overlap with human reviews over correctness. However, since human reviews…
DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding
Hengchuan Zhu, Yihuan Xu, Yichen Li +2
Recent advances in large language models (LLMs) and medical LLMs (Med-LLMs) have demonstrated strong performance on general medical benchmarks. However, their capabilities in speci…
FairSteer: Inference Time Debiasing for LLMs with Dynamic Activation Steering
Yichen Li, Zhiting Fan, Ruizhe Chen +4
Large language models (LLMs) are prone to capturing biases from training corpus, leading to potential negative social impacts. Existing prompt-based debiasing methods exhibit insta…
Identifying and Mitigating Social Bias Knowledge in Language Models
Ruizhe Chen, Yichen Li, Jianfei Yang +3
Generating fair and accurate predictions plays a pivotal role in deploying large language models (LLMs) in the real world. However, existing debiasing methods inevitably generate u…