activity
20162024
most citedTreeView: Peeking into Deep Neural Networks Via Feature-Space Partitioning

45 citations · 52 across the 11 of their papers we have counts for

collaborators

14 papers

cs.CL2024

Granite Guardian

Inkit Padhi, Manish Nagireddy, Giandomenico Cornacchia +20

We introduce the Granite Guardian models, a suite of safeguards designed to provide risk detection for prompts and responses, enabling safe and responsible use in combination with…

cs.CL2024

Graph-based Uncertainty Metrics for Long-form Language Model Outputs

Mingjian Jiang, Yangjun Ruan, Prasanna Sattigeri +2

Recent advancements in Large Language Models (LLMs) have significantly improved text generation capabilities, but these systems are still known to hallucinate, and granular uncerta…

cs.CL2024

Value Alignment from Unstructured Text

Inkit Padhi, Karthikeyan Natesan Ramamurthy, Prasanna Sattigeri +3

Aligning large language models (LLMs) to value systems has emerged as a significant area of research within the fields of AI and NLP. Currently, this alignment process relies on th…

cs.CL20241 cited

WikiContradict: A Benchmark for Evaluating LLMs on Real-World Knowledge Conflicts from Wikipedia

Yufang Hou, Alessandra Pascale, Javier Carnerero-Cano +5

Retrieval-augmented generation (RAG) has emerged as a promising solution to mitigate the limitations of large language models (LLMs), such as hallucinations and outdated informatio…

cs.AI20241 cited

Contextual Moral Value Alignment Through Context-Based Aggregation

Pierre Dognin, Jesus Rios, Ronny Luss +7

Developing value-aligned AI agents is a complex undertaking and an ongoing challenge in the field of AI. Specifically within the domain of Large Language Models (LLMs), the capabil…

cs.CL2024

Alignment Studio: Aligning Large Language Models to Particular Contextual Regulations

Swapnaja Achintalwar, Ioana Baldini, Djallel Bouneffouf +16

The alignment of large language models is usually done by model providers to add or control behaviors that are common or universally understood across use cases and contexts. In co…