collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

TimeSage-MT: A Multi-Turn Benchmark for Evaluating Agentic Time Series Reasoning

Yaxuan Kong, Qingren Yao, Yuqi Nie +7

Time series data inform critical decisions across many real-world domains. While large language model (LLM) agents can analyze data through natural language and tools, it remains u…

cs.CL2026

CoCoReviewBench: A Completeness- and Correctness-Oriented Benchmark for AI Reviewers

Hexuan Deng, Xiaopeng Ke, Yichen Li +6

Despite the rapid development of AI reviewers, evaluating such systems remains challenging: metrics favor overlap with human reviews over correctness. However, since human reviews…

cs.CL2025

DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding

Hengchuan Zhu, Yihuan Xu, Yichen Li +2

Recent advances in large language models (LLMs) and medical LLMs (Med-LLMs) have demonstrated strong performance on general medical benchmarks. However, their capabilities in speci…

cs.CL2025

FairSteer: Inference Time Debiasing for LLMs with Dynamic Activation Steering

Yichen Li, Zhiting Fan, Ruizhe Chen +4

Large language models (LLMs) are prone to capturing biases from training corpus, leading to potential negative social impacts. Existing prompt-based debiasing methods exhibit insta…

cs.CL2025

Identifying and Mitigating Social Bias Knowledge in Language Models

Ruizhe Chen, Yichen Li, Jianfei Yang +3

Generating fair and accurate predictions plays a pivotal role in deploying large language models (LLMs) in the real world. However, existing debiasing methods inevitably generate u…