31 citations · 45 across the 12 of their papers we have counts for
11 papers · 1 filter
Granite Guardian
Inkit Padhi, Manish Nagireddy, Giandomenico Cornacchia +20
We introduce the Granite Guardian models, a suite of safeguards designed to provide risk detection for prompts and responses, enabling safe and responsible use in combination with…
Value Alignment from Unstructured Text
Inkit Padhi, Karthikeyan Natesan Ramamurthy, Prasanna Sattigeri +3
Aligning large language models (LLMs) to value systems has emerged as a significant area of research within the fields of AI and NLP. Currently, this alignment process relies on th…
When in Doubt, Cascade: Towards Building Efficient and Capable Guardrails
Manish Nagireddy, Inkit Padhi, Soumya Ghosh +1
Large language models (LLMs) have convincing performance in a variety of downstream tasks. However, these systems are prone to generating undesirable outputs such as harmful and bi…
WikiContradict: A Benchmark for Evaluating LLMs on Real-World Knowledge Conflicts from Wikipedia
Yufang Hou, Alessandra Pascale, Javier Carnerero-Cano +5
Retrieval-augmented generation (RAG) has emerged as a promising solution to mitigate the limitations of large language models (LLMs), such as hallucinations and outdated informatio…
Alignment Studio: Aligning Large Language Models to Particular Contextual Regulations
Swapnaja Achintalwar, Ioana Baldini, Djallel Bouneffouf +16
The alignment of large language models is usually done by model providers to add or control behaviors that are common or universally understood across use cases and contexts. In co…
ReGen: Reinforcement Learning for Text and Knowledge Base Generation using Pretrained Language Models
Pierre L. Dognin, Inkit Padhi, Igor Melnyk +1
Automatic construction of relevant Knowledge Bases (KBs) from text, and generation of semantically meaningful text from KBs are both long-standing goals in Machine Learning. In thi…