1 citations · 1 across the 3 of their papers we have counts for
8 papers
DigitalCoach: Communication and Grounding Gaps in Human and Agentic Computer Use Coaching
Meng Chen, Anya Ji, Tsung-Han Wu +4
Agents are increasingly capable of automating software tasks, but can they teach humans how to use software themselves? We introduce DigitalCoach, a multimodal dataset of 72 human…
Are Large Reasoning Models Interruptible?
Tsung-Han Wu, Mihran Miroyan, David M. Chan +3
Real-world applications of Large Reasoning Models (LRMs) often require reasoning about changing prompts or environments. In this work, we challenge the frozen world assumption and…
Search Arena: Analyzing Search-Augmented LLMs
Mihran Miroyan, Tsung-Han Wu, Logan King +8
Search-augmented language models combine web search with Large Language Models (LLMs) to improve response groundedness and freshness. However, analyzing these systems remains chall…
Generate, but Verify: Reducing Hallucination in Vision-Language Models with Retrospective Resampling
Tsung-Han Wu, Heekyung Lee, Jiaxin Ge +3
Vision-Language Models (VLMs) excel at visual understanding but often suffer from visual hallucinations, where they generate descriptions of nonexistent objects, actions, or concep…
Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint
Heekyung Lee, Jiaxin Ge, Tsung-Han Wu +3
Rebus puzzles, visual riddles that encode language through imagery, spatial arrangement, and symbolic substitution, pose a unique challenge to current vision-language models (VLMs)…
CLAIR-A: Leveraging Large Language Models to Judge Audio Captions
Tsung-Han Wu, Joseph E. Gonzalez, Trevor Darrell +1
The Automated Audio Captioning (AAC) task asks models to generate natural language descriptions of an audio input. Evaluating these machine-generated audio captions is a complex ta…