20 papers
DigitalCoach: Communication and Grounding Gaps in Human and Agentic Computer Use Coaching
Meng Chen, Anya Ji, Tsung-Han Wu +4
Agents are increasingly capable of automating software tasks, but can they teach humans how to use software themselves? We introduce DigitalCoach, a multimodal dataset of 72 human…
Animation2Code: Evaluating Temporal Visual Reasoning in Video-to-Code Generation
Anya Ji, Abhijith Varma Mudunuri, David M. Chan +1
While recent vision-language models (VLMs) have achieved significant improvements on static visual-to-code tasks such as generating code for webpages, charts, or SVGs, it remains u…
T-Rex: Tactile-Reactive Dexterous Manipulation
Dantong Niu, Zhuoyang Liu, Zekai Wang +31
The ability to react dynamically to tactile signals has long been considered crucial to agile human-level dexterity. Yet contemporary learning-based Vision-Language-Action (VLA) mo…
Playful Agentic Robot Learning
Junyi Zhang, Jiaxin Ge, Hanjun Yoo +17
Current agentic robot systems can write executable Code-as-Policy programs, observe feedback, and revise behavior across multiple attempts, but they remain largely task-driven: reu…
Unintended Effects of Geographic Conditioning in Large Language Models
Naz Col, David M. Chan
Modern conversational AI systems frequently rely on user metadata to localize responses, yet the unintended regional biases introduced by this hidden context remain poorly understo…
Stateful Visual Encoders for Vision-Language Models
Zirui Wang, Junwei Yu, Adam Yala +3
Vision-language models (VLMs) are increasingly used in multi-image, multi-turn agentic settings where decisions depend on visual changes. However, in existing open-weight VLMs, vis…