11 papers
Rank4Gen: RAG-Preference-Aligned Document Set Selection and Ranking
Yongqi Fan, Yuxiang Chu, Zhentao Xia +9
In the RAG paradigm, document ranking determines the evidence available to downstream generators. Through controlled analysis, we identify two phenomena underexplored by existing r…
ClinicalMC: A Benchmark for Multi-Course Clinical Decision-Making with Large Language Models
Ruihui Hou, Siyi Zhu, Ziyue Huai +4
Large language models (LLMs) have been widely adopted in healthcare, yet they still encounter significant challenges in complex clinical decision-making scenarios. Existing benchma…
PsychÄChat: An Empathic Framework Focused on Emotion Shift Tracking and Safety Risk Analysis in Psychological Counseling
Zhentao Xia, Yongqi Fan, Yuxiang Chu +4
Large language models (LLMs) have demonstrated notable advancements in psychological counseling. However, existing models generally do not explicitly model seekers' emotion shifts…
TFRank: Think-Free Reasoning Enables Practical Pointwise LLM Ranking
Yongqi Fan, Xiaoyang Chen, Dezhi Ye +6
Reasoning-intensive ranking models built on Large Language Models (LLMs) have made notable progress. However, existing approaches often rely on large-scale LLMs and explicit Chain-…
KG-o1: Enhancing Multi-hop Question Answering in Large Language Models via Knowledge Graph Integration
Nan Wang, Yongqi Fan, yansha zhu +6
Large Language Models (LLMs) face challenges in knowledge-intensive reasoning tasks like classic multi-hop question and answering, which involves reasoning across multiple facts. T…
CMQCIC-Bench: A Chinese Benchmark for Evaluating Large Language Models in Medical Quality Control Indicator Calculation
Guangya Yu, Yanhao Li, Zongying Jiang +9
Medical quality control indicators are essential to assess the qualifications of healthcare institutions for medical services. With the impressive performance of large language mod…