Publications (64)
A Hierarchical Recurrent Encoder-Decoder For Generative Context-Aware Query Suggestion
Alessandro Sordoni, Yoshua Bengio, Hossein Vahabi +3
V-STaR: Training Verifiers for Self-Taught Reasoners
Arian Hosseini, Xingdi Yuan, Nikolay Malkin +3
Using Representation Expressiveness and Learnability to Evaluate Self-Supervised Learning Methods
Yuchen Lu, Zhen Liu, Aristide Baratin +3
Self-training with Few-shot Rationalization: Teacher Explanations Aid Student in Few-shot NLU
Meghana Moorthy Bhat, Alessandro Sordoni, Subhabrata Mukherjee
Augmented CycleGAN: Learning Many-to-Many Mappings from Unpaired Data
Amjad Almahairi, Sai Rajeswar, Alessandro Sordoni +2
Ordered Memory
Yikang Shen, Shawn Tan, Arian Hosseini +3
Twin Networks: Matching the Future for Sequence Generation
Dmitriy Serdyuk, Nan Rosemary Ke, Alessandro Sordoni +3
Decomposed Mutual Information Estimation for Contrastive Representation Learning
Alessandro Sordoni, Nouha Dziri, Hannes Schulz +3
Learning to Extract Context for Context-Aware LLM Inference
Minseon Kim, Lucas Caccia, Zhengyan Shi +4
Iterative Alternating Neural Attention for Machine Reading
Alessandro Sordoni, Philip Bachman, Adam Trischler +1
Multi-Head Adapter Routing for Cross-Task Generalization
Lucas Caccia, Edoardo Ponti, Zhan Su +3
The Emergence of the Shape Bias Results from Communicative Efficiency
Eva Portelance, Michael C. Frank, Dan Jurafsky +2
Test-Time Learning with an Evolving Library
Weijia Xu, Alessandro Sordoni, Chandan Singh +4
The paper introduces EvoLib, a test-time learning framework that lets large language models build, reuse, and evolve a shared library of knowledge abstractions across tasks without…
Not All LLM Reasoners Are Created Equal
Arian Hosseini, Alessandro Sordoni, Daniel Toyama +2
Exploring and Predicting Transferability across NLP Tasks
Tu Vu, Tong Wang, Tsendsuren Munkhdalai +5
An Empirical Study of Example Forgetting during Deep Neural Network Learning
Mariya Toneva, Alessandro Sordoni, Remi Tachet des Combes +3
deltaBLEU: A Discriminative Metric for Generation Tasks with Intrinsically Diverse Targets
Michel Galley, Chris Brockett, Alessandro Sordoni +6
Exploring Sparse Adapters for Scalable Merging of Parameter Efficient Experts
Samin Yeasar Arnob, Zhan Su, Minseon Kim +6
debug-gym: A Text-Based Environment for Interactive Debugging
Xingdi Yuan, Morgane M Moss, Charbel El Feghali +8
Looking at Vector Space and Language Models for IR using Density Matrices
Alessandro Sordoni, Jian-Yun Nie
Towards Information-Seeking Agents
Philip Bachman, Alessandro Sordoni, Adam Trischler
Does Pre-training Induce Systematic Inference? How Masked Language Models Acquire Commonsense Knowledge
Ian Porada, Alessandro Sordoni, Jackie Chi Kit Cheung
Understanding by Understanding Not: Modeling Negation in Language Models
Arian Hosseini, Siva Reddy, Dzmitry Bahdanau +3
Joint Prompt Optimization of Stacked LLMs using Variational Inference
Alessandro Sordoni, Xingdi Yuan, Marc-Alexandre Côté +6
On the Compositional Generalization Gap of In-Context Learning
Arian Hosseini, Ankit Vani, Dzmitry Bahdanau +2
Building End-To-End Dialogue Systems Using Generative Hierarchical Neural Network Models
Iulian V. Serban, Alessandro Sordoni, Yoshua Bengio +2
Ordered Neurons: Integrating Tree Structures into Recurrent Neural Networks
Yikang Shen, Shawn Tan, Alessandro Sordoni +1
Putting the Value Back in RL: Better Test-Time Scaling by Unifying LLM Reasoners With Verifiers
Kusha Sareen, Morgane M Moss, Alessandro Sordoni +2
Hybrid Generative-Retrieval Transformers for Dialogue Domain Adaptation
Igor Shalyminov, Alessandro Sordoni, Adam Atkinson +1
Straight to the Tree: Constituency Parsing with Neural Syntactic Distance
Yikang Shen, Zhouhan Lin, Athul Paul Jacob +3
Orchard: An Open-Source Agentic Modeling Framework
Baolin Peng, Wenlin Yao, Qianhui Wu +11
Gistify! Codebase-Level Understanding via Runtime Execution
Hyunji Lee, Minseon Kim, Chinmay Singh +10
Explicitly Modeling Syntax in Language Models with Incremental Parsing and a Dynamic Oracle
Yikang Shen, Shawn Tan, Alessandro Sordoni +2
A Modular Approach for Clinical SLMs Driven by Synthetic Data with Pre-Instruction Tuning, Model Merging, and Clinical-Tasks Alignment
Jean-Philippe Corbeil, Amin Dada, Jean-Michel Attendu +7
Better Language Model with Hypernym Class Prediction
He Bai, Tong Wang, Alessandro Sordoni +1
BugPilot: Complex Bug Generation for Efficient Learning of SWE Skills
Atharv Sonwane, Isadora White, Hyunji Lee +8
The Markovian Thinker: Architecture-Agnostic Linear Scaling of Reasoning
Milad Aghajohari, Kamran Chitsaz, Amirhossein Kazemnejad +4
Z-Forcing: Training Stochastic Recurrent Networks
Anirudh Goyal, Alessandro Sordoni, Marc-Alexandre Côté +2
Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
Gabriele Prato, Shagun Sodhani, Alessandro Sordoni +1
A Hierarchical Latent Variable Encoder-Decoder Model for Generating Dialogues
Iulian Vlad Serban, Alessandro Sordoni, Ryan Lowe +4
Linguistic Dependencies and Statistical Dependence
Jacob Louis Hoover, Alessandro Sordoni, Wenyu Du +1
A Survey on Model MoErging: Recycling and Routing Among Specialized Experts for Collaborative Learning
Prateek Yadav, Colin Raffel, Mohammed Muqeeth +6
Learning Algorithms for Active Learning
Philip Bachman, Alessandro Sordoni, Adam Trischler
VFunc: a Deep Generative Model for Functions
Philip Bachman, Riashat Islam, Alessandro Sordoni +1
Evaluating Distributional Distortion in Neural Language Modeling
Benjamin LeBrun, Alessandro Sordoni, Timothy J. O'Donnell
Improving Context-Aware Preference Modeling for Language Models
Silviu Pitis, Ziang Xiao, Nicolas Le Roux +1
VinePPO: Refining Credit Assignment in RL Training of LLMs
Amirhossein Kazemnejad, Milad Aghajohari, Eva Portelance +4
Towards Modular LLMs by Building and Reusing a Library of LoRAs
Oleksiy Ostapenko, Zhan Su, Edoardo Maria Ponti +5
Learning to Solve Complex Problems via Dataset Decomposition
Wanru Zhao, Lucas Caccia, Zhengyan Shi +3
Recursive Top-Down Production for Sentence Generation with Latent Trees
Shawn Tan, Yikang Shen, Timothy J. O'Donnell +2
A Neural Network Approach to Context-Sensitive Generation of Conversational Responses
Alessandro Sordoni, Michel Galley, Michael Auli +6
Trade-offs in Ensembling, Merging and Routing Among Parameter-Efficient Experts
Sanae Lotfi, Lucas Caccia, Alessandro Sordoni +2
Guiding Language Model Reasoning with Planning Tokens
Xinyi Wang, Lucas Caccia, Oleksiy Ostapenko +3
NewsQA: A Machine Comprehension Dataset
Adam Trischler, Tong Wang, Xingdi Yuan +4
Efficient Adversarial Training in LLMs with Continuous Attacks
Sophie Xhonneux, Alessandro Sordoni, Stephan Günnemann +2
Training Plug-n-Play Knowledge Modules with Deep Context Distillation
Lucas Caccia, Alan Ansell, Edoardo Ponti +2
MedRiskEval: Medical Risk Evaluation Benchmark of Language Models, On the Importance of User Perspectives in Healthcare Settings
Jean-Philippe Corbeil, Minseon Kim, Maxime Griot +4
Report on the First HIPstIR Workshop on the Future of Information Retrieval
Laura Dietz, Bhaskar Mitra, Jeremy Pickens +21
Counting to Explore and Generalize in Text-based Games
Xingdi Yuan, Marc-Alexandre Côté, Alessandro Sordoni +4
Combining Modular Skills in Multitask Learning
Edoardo M. Ponti, Alessandro Sordoni, Yoshua Bengio +1
Increasing Robustness to Spurious Correlations using Forgettable Examples
Yadollah Yaghoobzadeh, Soroush Mehri, Remi Tachet +2
Machine Comprehension by Text-to-Text Neural Question Generation
Xingdi Yuan, Tong Wang, Caglar Gulcehre +5
Metalearned Neural Memory
Tsendsuren Munkhdalai, Alessandro Sordoni, Tong Wang +1
Focused Hierarchical RNNs for Conditional Sequence Processing
Nan Rosemary Ke, Konrad Zolna, Alessandro Sordoni +6