4 papers
Steered Generation via Gradient-Based Optimization on Sparse Query Features
Sumanta Bhattacharyya, Pedram Rooshenas
Latent steering exploits internal representations of Large Language Models (LLMs) to guide generation, yet interventions on dense states can entangle distinct semantic features. In…
Generating Place-Based Compromises Between Two Points of View
Sumanta Bhattacharyya, Francine Chen, Scott Carter +6
Large Language Models (LLMs) excel academically but struggle with social intelligence tasks, such as creating good compromises. In this paper, we present methods for generating emp…
Audited Reasoning Refinement: Fine-Tuning Language Models via LLM-Guided Step-Wise Evaluation and Correction
Sumanta Bhattacharyya, Sara Riazi, Pedram Rooshenas
Training a task-specific small reasoning model is challenging when direct human supervision or high-quality labels are scarce. However, LLMs with reasoning capabilities produce abu…
Steered Generation via Gradient Descent on Sparse Features
Sumanta Bhattacharyya, Pedram Rooshenas
Large language models (LLMs) encode a diverse range of linguistic features within their latent representations, which can be harnessed to steer their output toward specific target…