3 citations · 7 across the 8 of their papers we have counts for
18 papers
Gained in Translation: Privileged Pairwise Judges Enhance Multilingual Reasoning
Lintang Sutawika, Gokul Swamy, Zhiwei Steven Wu +1
When asked a question in a language less seen in its training data, current reasoning large language models (RLMs) often exhibit dramatically lower performance than when asked the…
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
Charlie Zhang, Graham Neubig, Xiang Yue
Recent reinforcement learning (RL) techniques have yielded impressive reasoning improvements in language models, yet it remains unclear whether post-training truly extends a model'…
RefineBench: Evaluating Refinement Capability of Language Models via Checklists
Young-Jun Lee, Seungone Kim, Byung-Kwan Lee +6
Can language models (LMs) self-refine their own responses? This question is increasingly relevant as a wide range of real-world user interactions involve refinement requests. Howev…
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
André G. Viveiros, Patrick Fernandes, Saul Santos +7
Despite significant advances in vision-language models (VLMs), most existing work follows an English-centric design process, limiting their effectiveness in multilingual settings.…
Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities
Amanda Bertsch, Adithya Pratapa, Teruko Mitamura +2
As model context lengths continue to grow, concerns about whether models effectively use the full context length have persisted. While several carefully designed long-context evalu…
Accumulating Context Changes the Beliefs of Language Models
Jiayi Geng, Howard Chen, Ryan Liu +4
Language model (LM) assistants are increasingly used in applications such as brainstorming and research. Improvements in memory and context size have allowed these models to become…