4 papers · 1 filter
AI Steerability 360: A Toolkit for Steering Large Language Models
Erik Miehling, Karthikeyan Natesan Ramamurthy, Praveen Venkateswaran +10
The AI Steerability 360 toolkit is an extensible, open-source Python library for steering LLMs. Steering abstractions are designed around four model control surfaces: input (modifi…
Ranking Large Language Models without Ground Truth
Amit Dhurandhar, Rahul Nair, Moninder Singh +2
Evaluation and ranking of large language models (LLMs) has become an important problem with the proliferation of these models and their impact. Evaluation methods either require hu…
Reasoning about concepts with LLMs: Inconsistencies abound
Rosario Uceda-Sosa, Karthikeyan Natesan Ramamurthy, Maria Chang +1
The ability to summarize and organize knowledge into abstract concepts is key to learning and reasoning. Many industrial applications rely on the consistent and systematic use of c…
Alignment Studio: Aligning Large Language Models to Particular Contextual Regulations
Swapnaja Achintalwar, Ioana Baldini, Djallel Bouneffouf +16
The alignment of large language models is usually done by model providers to add or control behaviors that are common or universally understood across use cases and contexts. In co…