NewEvery arXiv paper, its researchers & institutions — mapped.
papers

Publications (64)

cs.NE2015

A Hierarchical Recurrent Encoder-Decoder For Generative Context-Aware Query Suggestion

Alessandro Sordoni, Yoshua Bengio, Hossein Vahabi +3

cs.LG2024

V-STaR: Training Verifiers for Self-Taught Reasoners

Arian Hosseini, Xingdi Yuan, Nikolay Malkin +3

cs.LG2023

Using Representation Expressiveness and Learnability to Evaluate Self-Supervised Learning Methods

Yuchen Lu, Zhen Liu, Aristide Baratin +3

cs.CL2021

Self-training with Few-shot Rationalization: Teacher Explanations Aid Student in Few-shot NLU

Meghana Moorthy Bhat, Alessandro Sordoni, Subhabrata Mukherjee

cs.LG2018

Augmented CycleGAN: Learning Many-to-Many Mappings from Unpaired Data

Amjad Almahairi, Sai Rajeswar, Alessandro Sordoni +2

cs.LG2019

Ordered Memory

Yikang Shen, Shawn Tan, Arian Hosseini +3

cs.LG2018

Twin Networks: Matching the Future for Sequence Generation

Dmitriy Serdyuk, Nan Rosemary Ke, Alessandro Sordoni +3

cs.LG2021

Decomposed Mutual Information Estimation for Contrastive Representation Learning

Alessandro Sordoni, Nouha Dziri, Hannes Schulz +3

cs.LG2025

Learning to Extract Context for Context-Aware LLM Inference

Minseon Kim, Lucas Caccia, Zhengyan Shi +4

cs.CL2016

Iterative Alternating Neural Attention for Machine Reading

Alessandro Sordoni, Philip Bachman, Adam Trischler +1

cs.AI2023

Multi-Head Adapter Routing for Cross-Task Generalization

Lucas Caccia, Edoardo Ponti, Zhan Su +3

cs.CL2021

The Emergence of the Shape Bias Results from Communicative Efficiency

Eva Portelance, Michael C. Frank, Dan Jurafsky +2

cs.LG2026

Test-Time Learning with an Evolving Library

Weijia Xu, Alessandro Sordoni, Chandan Singh +4

The paper introduces EvoLib, a test-time learning framework that lets large language models build, reuse, and evolve a shared library of knowledge abstractions across tasks without…

#test-time learning#large language models#knowledge library#continual learning
cs.LG2024

Not All LLM Reasoners Are Created Equal

Arian Hosseini, Alessandro Sordoni, Daniel Toyama +2

cs.CL2020

Exploring and Predicting Transferability across NLP Tasks

Tu Vu, Tong Wang, Tsendsuren Munkhdalai +5

cs.LG2019

An Empirical Study of Example Forgetting during Deep Neural Network Learning

Mariya Toneva, Alessandro Sordoni, Remi Tachet des Combes +3

cs.CL2015

deltaBLEU: A Discriminative Metric for Generation Tasks with Intrinsically Diverse Targets

Michel Galley, Chris Brockett, Alessandro Sordoni +6

cs.LG2025

Exploring Sparse Adapters for Scalable Merging of Parameter Efficient Experts

Samin Yeasar Arnob, Zhan Su, Minseon Kim +6

cs.AI2025

debug-gym: A Text-Based Environment for Interactive Debugging

Xingdi Yuan, Morgane M Moss, Charbel El Feghali +8

cs.IR2014

Looking at Vector Space and Language Models for IR using Density Matrices

Alessandro Sordoni, Jian-Yun Nie

cs.LG2016

Towards Information-Seeking Agents

Philip Bachman, Alessandro Sordoni, Adam Trischler

cs.CL2021

Does Pre-training Induce Systematic Inference? How Masked Language Models Acquire Commonsense Knowledge

Ian Porada, Alessandro Sordoni, Jackie Chi Kit Cheung

cs.CL2021

Understanding by Understanding Not: Modeling Negation in Language Models

Arian Hosseini, Siva Reddy, Dzmitry Bahdanau +3

cs.CL2023

Joint Prompt Optimization of Stacked LLMs using Variational Inference

Alessandro Sordoni, Xingdi Yuan, Marc-Alexandre Côté +6

cs.CL2022

On the Compositional Generalization Gap of In-Context Learning

Arian Hosseini, Ankit Vani, Dzmitry Bahdanau +2

cs.CL2016

Building End-To-End Dialogue Systems Using Generative Hierarchical Neural Network Models

Iulian V. Serban, Alessandro Sordoni, Yoshua Bengio +2

cs.CL2019

Ordered Neurons: Integrating Tree Structures into Recurrent Neural Networks

Yikang Shen, Shawn Tan, Alessandro Sordoni +1

cs.LG2026

Putting the Value Back in RL: Better Test-Time Scaling by Unifying LLM Reasoners With Verifiers

Kusha Sareen, Morgane M Moss, Alessandro Sordoni +2

cs.CL2020

Hybrid Generative-Retrieval Transformers for Dialogue Domain Adaptation

Igor Shalyminov, Alessandro Sordoni, Adam Atkinson +1

cs.CL2018

Straight to the Tree: Constituency Parsing with Neural Syntactic Distance

Yikang Shen, Zhouhan Lin, Athul Paul Jacob +3

cs.AI2026

Orchard: An Open-Source Agentic Modeling Framework

Baolin Peng, Wenlin Yao, Qianhui Wu +11

cs.CL2025

Gistify! Codebase-Level Understanding via Runtime Execution

Hyunji Lee, Minseon Kim, Chinmay Singh +10

cs.CL2021

Explicitly Modeling Syntax in Language Models with Incremental Parsing and a Dynamic Oracle

Yikang Shen, Shawn Tan, Alessandro Sordoni +2

cs.CL2025

A Modular Approach for Clinical SLMs Driven by Synthetic Data with Pre-Instruction Tuning, Model Merging, and Clinical-Tasks Alignment

Jean-Philippe Corbeil, Amin Dada, Jean-Michel Attendu +7

cs.CL2022

Better Language Model with Hypernym Class Prediction

He Bai, Tong Wang, Alessandro Sordoni +1

cs.SE2025

BugPilot: Complex Bug Generation for Efficient Learning of SWE Skills

Atharv Sonwane, Isadora White, Hyunji Lee +8

cs.LG2025

The Markovian Thinker: Architecture-Agnostic Linear Scaling of Reasoning

Milad Aghajohari, Kamran Chitsaz, Amirhossein Kazemnejad +4

stat.ML2017

Z-Forcing: Training Stochastic Recurrent Networks

Anirudh Goyal, Alessandro Sordoni, Marc-Alexandre Côté +2

cs.CL2025

Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models

Gabriele Prato, Shagun Sodhani, Alessandro Sordoni +1

cs.CL2016

A Hierarchical Latent Variable Encoder-Decoder Model for Generating Dialogues

Iulian Vlad Serban, Alessandro Sordoni, Ryan Lowe +4

cs.CL2022

Linguistic Dependencies and Statistical Dependence

Jacob Louis Hoover, Alessandro Sordoni, Wenyu Du +1

cs.LG2025

A Survey on Model MoErging: Recycling and Routing Among Specialized Experts for Collaborative Learning

Prateek Yadav, Colin Raffel, Mohammed Muqeeth +6

cs.LG2017

Learning Algorithms for Active Learning

Philip Bachman, Alessandro Sordoni, Adam Trischler

cs.LG2018

VFunc: a Deep Generative Model for Functions

Philip Bachman, Riashat Islam, Alessandro Sordoni +1

cs.CL2022

Evaluating Distributional Distortion in Neural Language Modeling

Benjamin LeBrun, Alessandro Sordoni, Timothy J. O'Donnell

cs.CL2024

Improving Context-Aware Preference Modeling for Language Models

Silviu Pitis, Ziang Xiao, Nicolas Le Roux +1

cs.LG2025

VinePPO: Refining Credit Assignment in RL Training of LLMs

Amirhossein Kazemnejad, Milad Aghajohari, Eva Portelance +4

cs.LG2024

Towards Modular LLMs by Building and Reusing a Library of LoRAs

Oleksiy Ostapenko, Zhan Su, Edoardo Maria Ponti +5

cs.LG2026

Learning to Solve Complex Problems via Dataset Decomposition

Wanru Zhao, Lucas Caccia, Zhengyan Shi +3

cs.CL2020

Recursive Top-Down Production for Sentence Generation with Latent Trees

Shawn Tan, Yikang Shen, Timothy J. O'Donnell +2

cs.CL2015

A Neural Network Approach to Context-Sensitive Generation of Conversational Responses

Alessandro Sordoni, Michel Galley, Michael Auli +6

cs.LG2026

Trade-offs in Ensembling, Merging and Routing Among Parameter-Efficient Experts

Sanae Lotfi, Lucas Caccia, Alessandro Sordoni +2

cs.CL2024

Guiding Language Model Reasoning with Planning Tokens

Xinyi Wang, Lucas Caccia, Oleksiy Ostapenko +3

cs.CL2017

NewsQA: A Machine Comprehension Dataset

Adam Trischler, Tong Wang, Xingdi Yuan +4

cs.LG2024

Efficient Adversarial Training in LLMs with Continuous Attacks

Sophie Xhonneux, Alessandro Sordoni, Stephan Günnemann +2

cs.LG2025

Training Plug-n-Play Knowledge Modules with Deep Context Distillation

Lucas Caccia, Alan Ansell, Edoardo Ponti +2

cs.CL2026

MedRiskEval: Medical Risk Evaluation Benchmark of Language Models, On the Importance of User Perspectives in Healthcare Settings

Jean-Philippe Corbeil, Minseon Kim, Maxime Griot +4

cs.IR2019

Report on the First HIPstIR Workshop on the Future of Information Retrieval

Laura Dietz, Bhaskar Mitra, Jeremy Pickens +21

cs.CL2019

Counting to Explore and Generalize in Text-based Games

Xingdi Yuan, Marc-Alexandre Côté, Alessandro Sordoni +4

cs.LG2022

Combining Modular Skills in Multitask Learning

Edoardo M. Ponti, Alessandro Sordoni, Yoshua Bengio +1

cs.CL2021

Increasing Robustness to Spurious Correlations using Forgettable Examples

Yadollah Yaghoobzadeh, Soroush Mehri, Remi Tachet +2

cs.CL2017

Machine Comprehension by Text-to-Text Neural Question Generation

Xingdi Yuan, Tong Wang, Caglar Gulcehre +5

cs.NE2019

Metalearned Neural Memory

Tsendsuren Munkhdalai, Alessandro Sordoni, Tong Wang +1

stat.ML2018

Focused Hierarchical RNNs for Conditional Sequence Processing

Nan Rosemary Ke, Konrad Zolna, Alessandro Sordoni +6