6 papers · 1 filter
DigitalCoach: Communication and Grounding Gaps in Human and Agentic Computer Use Coaching
Meng Chen, Anya Ji, Tsung-Han Wu +4
Agents are increasingly capable of automating software tasks, but can they teach humans how to use software themselves? We introduce DigitalCoach, a multimodal dataset of 72 human…
Unintended Effects of Geographic Conditioning in Large Language Models
Naz Col, David M. Chan
Modern conversational AI systems frequently rely on user metadata to localize responses, yet the unintended regional biases introduced by this hidden context remain poorly understo…
Are Large Reasoning Models Interruptible?
Tsung-Han Wu, Mihran Miroyan, David M. Chan +3
Real-world applications of Large Reasoning Models (LRMs) often require reasoning about changing prompts or environments. In this work, we challenge the frozen world assumption and…
Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint
Heekyung Lee, Jiaxin Ge, Tsung-Han Wu +3
Rebus puzzles, visual riddles that encode language through imagery, spatial arrangement, and symbolic substitution, pose a unique challenge to current vision-language models (VLMs)…
CLAIR-A: Leveraging Large Language Models to Judge Audio Captions
Tsung-Han Wu, Joseph E. Gonzalez, Trevor Darrell +1
The Automated Audio Captioning (AAC) task asks models to generate natural language descriptions of an audio input. Evaluating these machine-generated audio captions is a complex ta…
Enough Coin Flips Can Make LLMs Act Bayesian
Ritwik Gupta, Rodolfo Corona, Jiaxin Ge +4
Large language models (LLMs) exhibit the ability to generalize given few-shot examples in their input prompt, an emergent capability known as in-context learning (ICL). We investig…