From the 1 of 29 linked papers with an AI index.
26 papers · 1 filter
On the Limits of Model Merging for Multilinguality in Pre-Training
Seth Aycock, Fedor Vitiugin, Aleksandr Umnov +2
Endowing models with consistent multilingual performance can be achieved by mixing pre-training data, or post-training approaches such as language-specific model merging. In this w…
When Contextual Inference Fails: Cancelability in Interactive Instruction Following
Natalia Bila, Kata Naszádi, Kata Naszádi +2
We investigate the separation of literal interpretation from contextual inference in a collaborative block-building tasks, where an agent must resolve underspecified instructions u…
What Does Neuro Mean to Cardio? Investigating the Role of Clinical Specialty Data in Medical LLMs
Xinlan Yan, Di Wu, Yibin Lei +2
In this paper, we introduce S-MedQA, an English medical question-answering (QA) dataset designed for benchmarking large language models (LLMs) in fine-grained clinical specialties.…
Do Language Models Reason Across Languages?
Yan Meng, Wafaa Mohammed, Christof Monz
The real-world information sources are inherently multilingual, which naturally raises a question about whether language models can synthesize information across languages. In this…
Remedy-R: Generative Reasoning for Machine Translation Evaluation without Error Annotations
Shaomu Tan, Ryosuke Mitani, Ritvik Choudhary +3
Over the years, automatic MT metrics have hillclimbed benchmarks and presented strong and sometimes human-level agreement with human ratings. Yet they remain black-box, offering li…
Lost at the Beginning of Reasoning
Baohao Liao, Xinyi Chen, Sara Rajaee +5
Recent advancements in large language models (LLMs) have significantly advanced complex reasoning capabilities, particularly through extended chain-of-thought (CoT) reasoning that…