collaborators

6 papers

cs.AI2026

Tandem Reinforcement Learning with Verifiable Rewards

Difan Jiao, Raghav Singhal, Robert West +1

Reinforcement learning with verifiable rewards (RLVR) has significantly improved the reasoning capability of large language models, reaching expert or even superhuman performance i…

cs.LG2026

RL for Reasoning by Adaptively Revealing Rationales

Mohammad Hossein Amani, Aryo Lotfi, Nicolas Mario Baldwin +4

Learning in the combinatorially large output space of sequence generation problems is challenging as providing expert demonstrations scales poorly with sequence length, and RL stru…

cs.CL2026

Cmprsr: Abstractive Token-Level Question-Agnostic Prompt Compressor

Ivan Zakazov, Berke Argin, Oussama Gabouj +6

Motivated by the high costs of using black-box Large Language Models (LLMs), we introduce a novel prompt compression paradigm, under which we use smaller LLMs to compress inputs fo…

cs.CL2025

GRAD: Generative Retrieval-Aligned Demonstration Sampler for Efficient Few-Shot Reasoning

Oussama Gabouj, Kamel Charaf, Ivan Zakazov +2

Large Language Models (LLMs) achieve strong performance across diverse tasks, but their effectiveness often depends on the quality of the provided context. Retrieval-Augmented Gene…

cs.CY2025

Assessing Social Alignment: Do Personality-Prompted Large Language Models Behave Like Humans?

Ivan Zakazov, Mikolaj Boronski, Lorenzo Drudi +1

The ongoing revolution in language modeling has led to various novel applications, some of which rely on the emerging social abilities of large language models (LLMs). Already, man…

cs.CL2025

TRPrompt: Bootstrapping Query-Aware Prompt Optimization from Textual Rewards

Andreea Nica, Ivan Zakazov, Nicolas Mario Baldwin +2

Prompt optimization improves the reasoning abilities of large language models (LLMs) without requiring parameter updates to the target model. Following heuristic-based "Think step…