From the 1 of 4 linked papers with an AI index.
4 papers
Implicit Reasoning Steering via Concept Chaining
Xiao Ye, Sanika Chavan, Yuxi Huang +4
The paper introduces Concept Chaining, a method that creates short natural-language paragraphs linking question entities to a target answer via intermediate concepts, and uses cont…
Evaluating Medical LLMs by Levels of Autonomy: A Survey Moving from Benchmarks to Applications
Xiao Ye, Jacob Dineen, Zhaonan Li +11
Medical Large language models achieve strong scores on standard benchmarks; however, the transfer of those results to safe and reliable performance in clinical workflows remains a…
ArenaBencher: Automatic Benchmark Evolution via Multi-Model Competitive Evaluation
Qin Liu, Jacob Dineen, Yuxi Huang +4
Benchmarks are central to measuring the capabilities of large language models and guiding model development, yet widespread data leakage from pretraining corpora undermines their v…
Code-Survey: An LLM-Driven Methodology for Analyzing Large-Scale Codebases
Yusheng Zheng, Yiwei Yang, Haoqin Tu +1
Modern software systems like the Linux kernel are among the world's largest and most intricate codebases, continually evolving with new features and increasing complexity. Understa…