13 papers
Estimating Tail Risks in Language Model Output Distributions
Rico Angell, Raghav Singhal, Zachary Horvitz +4
Language models are increasingly capable and are being rapidly deployed on a population-level scale. As a result, the safety of these models is increasingly high-stakes. Fortunatel…
Solving Zebra Puzzles Using Constraint-Guided Multi-Agent Systems
Shmuel Berman, Kathleen McKeown, Baishakhi Ray
Prior research has enhanced the ability of Large Language Models (LLMs) to solve logic puzzles using techniques such as chain-of-thought prompting or introducing a symbolic represe…
Factuality on Demand: Controlling the Factuality-Informativeness Trade-off in Text Generation
Ziwei Gong, Yanda Chen, Julia Hirschberg +4
Large language models (LLMs) encode knowledge with varying degrees of confidence. When responding to queries, models face an inherent trade-off: they can generate responses that ar…
iBERT: Interpretable Embeddings via Sense Decomposition
Vishal Anand, Milad Alshomary, Kathleen McKeown
We present iBERT (interpretable-BERT), an encoder to produce inherently interpretable and controllable embeddings - designed to modularize and expose the discriminative cues presen…
Rethinking Reasoning with MDLMs: Early Exits, Post-hoc Reasoning, and Beyond
Zachary Horvitz, Raghav Singhal, Hao Zou +4
The reasoning paradigm, where language models reason before answering, has enabled breakthroughs on tasks such as mathematical problem-solving. While current tooling for reasoning…
Layered Insights: Generalizable Analysis of Authorial Style by Leveraging All Transformer Layers
Milad Alshomary, Nikhil Reddy Varimalla, Vishal Anand +2
We propose a new approach for the authorship attribution task that leverages the various linguistic representations learned at different layers of pre-trained transformer-based mod…