9 papers
DigitalCoach: Communication and Grounding Gaps in Human and Agentic Computer Use Coaching
Meng Chen, Anya Ji, Tsung-Han Wu +4
Agents are increasingly capable of automating software tasks, but can they teach humans how to use software themselves? We introduce DigitalCoach, a multimodal dataset of 72 human…
Animation2Code: Evaluating Temporal Visual Reasoning in Video-to-Code Generation
Anya Ji, Abhijith Varma Mudunuri, David M. Chan +1
While recent vision-language models (VLMs) have achieved significant improvements on static visual-to-code tasks such as generating code for webpages, charts, or SVGs, it remains u…
Representational Similarity and Model Behavior in Multi-Agent Interaction
Yujin Potter, Seun Eisape, Shiyang Lai +6
Researchers have shown that neural similarity among humans predicts social closeness and cooperative success, whereas innovation often emerges from interactions among dissimilar in…
ScribbleEdit: Synthetic Data for Image Editing with Scribbles and Text
Anya Ji, George Ma, Téa Wright +4
Recent progress in generative models has significantly advanced image editing capabilities, yet precise and intuitive user control remains difficult. Specifically, users often stru…
Long Chain-of-Thought Reasoning Across Languages
Josh Barua, Seun Eisape, Kayo Yin +1
While large reasoning models have shown remarkable ability to generate long chains-of-thought (CoTs) in English, we still lack understanding of how these long-form reasoning abilit…
Characterizing Language Use in a Collaborative Situated Game
Nicholas Tomlin, Naitian Zhou, Eve Fleisig +10
Cooperative video games, where multiple participants must coordinate by communicating and reasoning under uncertainty in complex environments, yield a rich source of language data.…