3 papers
cs.CL2026
Poor Alignment and Steerability of Large Language Models: Evidence from College Admission Essays
Jinsook Lee, AJ Alvero, Thorsten Joachims +1
People are increasingly using technologies equipped with large language models (LLM) to write texts for formal communication, which raises two important questions at the intersecti…
cs.LG2025
MultiScale Contextual Bandits for Long Term Objectives
Richa Rastogi, Yuta Saito, Thorsten Joachims
The feedback that AI systems (e.g., recommender systems, chatbots) collect from user interactions is a crucial source of training data. While short-term feedback (e.g., clicks, eng…
cs.LG2025
Prompt Optimization with Logged Bandit Data
Haruka Kiyohara, Daniel Yiming Cao, Yuta Saito +1
We study how to use naturally available user feedback, such as clicks, to optimize large language model (LLM) pipelines for generating personalized sentences using prompts. Naive a…