Publications (33)
CogMol: Target-Specific and Selective Drug Design for COVID-19 Using Deep Generative Models
Vijil Chenthamarakshan, Payel Das, Samuel C. Hoffman +8
The novel nature of SARS-CoV-2 calls for the development of efficient de novo drug design approaches. In this study, we propose an end-to-end framework, named CogMol (Controlled Ge…
Learning Implicit Text Generation via Feature Matching
Inkit Padhi, Pierre Dognin, Ke Bai +4
Generative feature matching network (GFMN) is an approach for training implicit generative models for images by performing moment matching on features from pre-trained neural netwo…
Answering the Wrong Question: Reasoning Trace Inversion for Abstention in LLMs
Abinitha Gourabathina, Inkit Padhi, Manish Nagireddy +2
For Large Language Models (LLMs) to be reliably deployed, models must effectively know when not to answer: abstain. Reasoning models, in particular, have gained attention for impre…
WikiContradict: A Benchmark for Evaluating LLMs on Real-World Knowledge Conflicts from Wikipedia
Yufang Hou, Alessandra Pascale, Javier Carnerero-Cano +5
Retrieval-augmented generation (RAG) has emerged as a promising solution to mitigate the limitations of large language models (LLMs), such as hallucinations and outdated informatio…
Large-Scale Chemical Language Representations Capture Molecular Structure and Properties
Jerret Ross, Brian Belgodere, Vijil Chenthamarakshan +3
Models based on machine learning can enable accurate and fast molecular property predictions, which is of interest in drug discovery and material design. Various supervised machine…
Sobolev Independence Criterion
Youssef Mroueh, Tom Sercu, Mattia Rigotti +2
We propose the Sobolev Independence Criterion (SIC), an interpretable dependency measure between a high dimensional random variable X and a response variable Y . SIC decomposes to…
PepCVAE: Semi-Supervised Targeted Design of Antimicrobial Peptide Sequences
Payel Das, Kahini Wadhawan, Oscar Chang +6
Given the emerging global threat of antimicrobial resistance, new methods for next-generation antimicrobial design are urgently needed. We report a peptide generation framework Pep…
The Impact of Positional Encoding on Length Generalization in Transformers
Amirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy +2
Length generalization, the ability to generalize from small training context sizes to larger ones, is a critical challenge in the development of Transformer-based language models.…
Granite Guardian
Inkit Padhi, Manish Nagireddy, Giandomenico Cornacchia +20
We introduce the Granite Guardian models, a suite of safeguards designed to provide risk detection for prompts and responses, enabling safe and responsible use in combination with…
Generate Your Counterfactuals: Towards Controlled Counterfactual Generation for Text
Nishtha Madaan, Inkit Padhi, Naveen Panwar +1
Machine Learning has seen tremendous growth recently, which has led to larger adoption of ML systems for educational assessments, credit risk, healthcare, employment, criminal just…
Value Alignment from Unstructured Text
Inkit Padhi, Karthikeyan Natesan Ramamurthy, Prasanna Sattigeri +3
Aligning large language models (LLMs) to value systems has emerged as a significant area of research within the fields of AI and NLP. Currently, this alignment process relies on th…
Reprogramming Pretrained Language Models for Antibody Sequence Infilling
Igor Melnyk, Vijil Chenthamarakshan, Pin-Yu Chen +4
Antibodies comprise the most versatile class of binding molecules, with numerous applications in biomedicine. Computational design of antibodies involves generating novel and diver…
Fair Infinitesimal Jackknife: Mitigating the Influence of Biased Training Data Points Without Refitting
Prasanna Sattigeri, Soumya Ghosh, Inkit Padhi +2
In consequential decision-making applications, mitigating unwanted biases in machine learning models that yield systematic disadvantage to members of groups delineated by sensitive…
Final-Model-Only Data Attribution with a Unifying View of Gradient-Based Methods
Dennis Wei, Inkit Padhi, Soumya Ghosh +3
Training data attribution (TDA) is concerned with understanding model behavior in terms of the training data. This paper draws attention to the common setting where one has access…
Building a Foundational Guardrail for General Agentic Systems via Synthetic Data
Yue Huang, Hang Hua, Yujun Zhou +11
While LLM agents can plan multi-step tasks, intervening at the planning stage-before any action is executed-is often the safest way to prevent harm, since certain risks can lead to…
Split, Unlearn, Merge: Leveraging Data Attributes for More Effective Unlearning in LLMs
Swanand Ravindra Kadhe, Farhan Ahmed, Dennis Wei +2
Large language models (LLMs) have shown to pose social and ethical risks such as generating toxic language or facilitating malicious use of hazardous knowledge. Machine unlearning…
Contextual Moral Value Alignment Through Context-Based Aggregation
Pierre Dognin, Jesus Rios, Ronny Luss +7
Developing value-aligned AI agents is a complex undertaking and an ongoing challenge in the field of AI. Specifically within the domain of Large Language Models (LLMs), the capabil…
Alignment Studio: Aligning Large Language Models to Particular Contextual Regulations
Swapnaja Achintalwar, Ioana Baldini, Djallel Bouneffouf +16
The alignment of large language models is usually done by model providers to add or control behaviors that are common or universally understood across use cases and contexts. In co…
DualTKB: A Dual Learning Bridge between Text and Knowledge Base
Pierre L. Dognin, Igor Melnyk, Inkit Padhi +2
In this work, we present a dual learning approach for unsupervised text to path and path to text transfers in Commonsense Knowledge Bases (KBs). We investigate the impact of weak s…
Learning Implicit Generative Models by Matching Perceptual Features
Cicero Nogueira dos Santos, Youssef Mroueh, Inkit Padhi +1
Perceptual features (PFs) have been used with great success in tasks such as transfer learning, style transfer, and super-resolution. However, the efficacy of PFs as key source of…
Tabular Transformers for Modeling Multivariate Time Series
Inkit Padhi, Yair Schiff, Igor Melnyk +6
Tabular datasets are ubiquitous in data science applications. Given their importance, it seems natural to apply state-of-the-art deep learning algorithms in order to fully unlock t…
Fighting Offensive Language on Social Media with Unsupervised Text Style Transfer
Cicero Nogueira dos Santos, Igor Melnyk, Inkit Padhi
We introduce a new approach to tackle the problem of offensive language in online social media. Our approach uses unsupervised text style transfer to translate offensive sentences…
Cloud-Based Real-Time Molecular Screening Platform with MolFormer
Brian Belgodere, Vijil Chenthamarakshan, Payel Das +9
With the prospect of automating a number of chemical tasks with high fidelity, chemical language processing models are emerging at a rapid speed. Here, we present a cloud-based rea…
Alleviating Noisy Data in Image Captioning with Cooperative Distillation
Pierre Dognin, Igor Melnyk, Youssef Mroueh +4
Image captioning systems have made substantial progress, largely due to the availability of curated datasets like Microsoft COCO or Vizwiz that have accurate descriptions of their…
Detectors for Safe and Reliable LLMs: Implementations, Uses, and Limitations
Swapnaja Achintalwar, Adriana Alvarado Garcia, Ateret Anaby-Tavor +35
Large language models (LLMs) are susceptible to a variety of risks, from non-faithful output to biased and toxic generations. Due to several limiting factors surrounding LLMs (trai…
Accelerating Material Design with the Generative Toolkit for Scientific Discovery
Matteo Manica, Jannis Born, Joris Cadow +21
With the growing availability of data within various scientific domains, generative models hold enormous potential to accelerate scientific discovery. They harness powerful represe…
Image Captioning as an Assistive Technology: Lessons Learned from VizWiz 2020 Challenge
Pierre Dognin, Igor Melnyk, Youssef Mroueh +6
Image captioning has recently demonstrated impressive progress largely owing to the introduction of neural network algorithms trained on curated dataset like MS-COCO. Often work in…
Improved Neural Text Attribute Transfer with Non-parallel Data
Igor Melnyk, Cicero Nogueira dos Santos, Kahini Wadhawan +2
Text attribute transfer using non-parallel data requires methods that can perform disentanglement of content and linguistic attributes. In this work, we propose multiple improvemen…
Auditing and Generating Synthetic Data with Controllable Trust Trade-offs
Brian Belgodere, Pierre Dognin, Adam Ivankay +11
Real-world data often exhibits bias, imbalance, and privacy risks. Synthetic datasets have emerged to address these issues. This paradigm relies on generative AI models to generate…
ReGen: Reinforcement Learning for Text and Knowledge Base Generation using Pretrained Language Models
Pierre L. Dognin, Inkit Padhi, Igor Melnyk +1
Automatic construction of relevant Knowledge Bases (KBs) from text, and generation of semantically meaningful text from KBs are both long-standing goals in Machine Learning. In thi…
When in Doubt, Cascade: Towards Building Efficient and Capable Guardrails
Manish Nagireddy, Inkit Padhi, Soumya Ghosh +1
Large language models (LLMs) have convincing performance in a variety of downstream tasks. However, these systems are prone to generating undesirable outputs such as harmful and bi…
Programming Refusal with Conditional Activation Steering
Bruce W. Lee, Inkit Padhi, Karthikeyan Natesan Ramamurthy +4
LLMs have shown remarkable capabilities, but precisely controlling their response behavior remains challenging. Existing activation steering methods alter LLM behavior indiscrimina…
Accelerating Antimicrobial Discovery with Controllable Deep Generative Models and Molecular Dynamics
Payel Das, Tom Sercu, Kahini Wadhawan +12
De novo therapeutic design is challenged by a vast chemical repertoire and multiple constraints, e.g., high broad-spectrum potency and low toxicity. We propose CLaSS (Controlled La…