activity
20182026
most citedGroup Equivariant Conditional Neural Processes

2 citations · 8 across the 74 of their papers we have counts for

collaborators
Showing cs.CLShow all

27 papers · 1 filter

cs.CL2026

HealMed: Multilingual Evaluation of Large Language Models in Medicine

Yingjian Chen, Fan Gao, Sherry T. Tong +42

We present HealMed, an expert-reviewed benchmark for multilingual evaluation of large language models in medicine. HealMed contains 1,000 examples in each of nine languages, drawn…

cs.CL2026

Batch-wise Adaptive Pruning: Periodic Neuron Activation-Aware Weight Pruning for Language Reasoning Model

Yongmin Kim, Shota Takashiro, Yusuke Iwasawa +2

Large Reasoning Models (LRMs) achieve strong performance on complex tasks through extended chain-of-thought generation, but incur substantial computational costs during inference.…

cs.CL2026

Bootstrapping Niche Multilingual Code Translation via Reinforcement Learning with Execution-Based Verifiable Supervision

Kouki Yuki, Jie Zeng, Kyoko Ogawa +6

Code translation must preserve executable behavior across many programming languages, yet neural code translation has largely focused on a few popular languages such as C++, Java,…

cs.CL2026

Clustered Self-Assessment: A Simple yet Effective Method for Uncertainty Quantification in Large Language Models

Qi Cao, Takeshi Kojima, Andrew Gambardella +3

Large language models (LLMs) demonstrate remarkable performance across diverse tasks, but they often generate responses that appear plausible while being factually incorrect. This…

cs.CL2026

Semantic Token Clustering for Efficient Uncertainty Quantification in Large Language Models

Qi Cao, Andrew Gambardella, Takeshi Kojima +2

Large language models (LLMs) have demonstrated remarkable capabilities across diverse tasks. However, the truthfulness of their outputs is not guaranteed, and their tendency toward…

cs.CL2026

Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models

Xiaojie Gu, Sherry T. Tong, Aosong Feng +8

Evaluating the reasoning abilities of large language models (LLMs) solely from final answers can obscure failures in intermediate steps, especially in multi-hop QA benchmarks witho…