2 citations · 8 across the 74 of their papers we have counts for
27 papers · 1 filter
HealMed: Multilingual Evaluation of Large Language Models in Medicine
Yingjian Chen, Fan Gao, Sherry T. Tong +42
We present HealMed, an expert-reviewed benchmark for multilingual evaluation of large language models in medicine. HealMed contains 1,000 examples in each of nine languages, drawn…
Batch-wise Adaptive Pruning: Periodic Neuron Activation-Aware Weight Pruning for Language Reasoning Model
Yongmin Kim, Shota Takashiro, Yusuke Iwasawa +2
Large Reasoning Models (LRMs) achieve strong performance on complex tasks through extended chain-of-thought generation, but incur substantial computational costs during inference.…
Bootstrapping Niche Multilingual Code Translation via Reinforcement Learning with Execution-Based Verifiable Supervision
Kouki Yuki, Jie Zeng, Kyoko Ogawa +6
Code translation must preserve executable behavior across many programming languages, yet neural code translation has largely focused on a few popular languages such as C++, Java,…
Clustered Self-Assessment: A Simple yet Effective Method for Uncertainty Quantification in Large Language Models
Qi Cao, Takeshi Kojima, Andrew Gambardella +3
Large language models (LLMs) demonstrate remarkable performance across diverse tasks, but they often generate responses that appear plausible while being factually incorrect. This…
Semantic Token Clustering for Efficient Uncertainty Quantification in Large Language Models
Qi Cao, Andrew Gambardella, Takeshi Kojima +2
Large language models (LLMs) have demonstrated remarkable capabilities across diverse tasks. However, the truthfulness of their outputs is not guaranteed, and their tendency toward…
Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models
Xiaojie Gu, Sherry T. Tong, Aosong Feng +8
Evaluating the reasoning abilities of large language models (LLMs) solely from final answers can obscure failures in intermediate steps, especially in multi-hop QA benchmarks witho…