1 citations · 1 across the 3 of their papers we have counts for
4 papers
RAPTOR: Ridge-Adaptive Logistic Probes
Ziqi Gao, Yaotian Zhu, Qingcheng Zeng +4
Probing studies what information is encoded in a frozen LLM's layer representations by training a lightweight predictor on top of them. Beyond analysis, probes are often used opera…
The Confidence Dichotomy: Analyzing and Mitigating Miscalibration in Tool-Use Agents
Weihao Xuan, Qingcheng Zeng, Heli Qi +3
Autonomous agents based on large language models (LLMs) are rapidly evolving to handle multi-turn tasks, but ensuring their trustworthiness remains a critical challenge. A fundamen…
Toward Global Large Language Models in Medicine
Rui Yang, Huitao Li, Weihao Xuan +47
Despite continuous advances in medical technology, the global distribution of health care resources remains uneven. The development of large language models (LLMs) has transformed…
MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation
Weihao Xuan, Rui Yang, Heli Qi +29
Existing large language model (LLM) evaluation benchmarks primarily focus on English, while current multilingual tasks lack parallel questions that specifically assess cross-lingui…