6 papers · 1 filter
CoVA-SFT: A Large-Scale Dataset for Chain of Visual Abstractions
Tsung-Han Wu, Heekyung Lee, Anya Ji +4
Chain-of-thought (CoT) reasoning has dramatically improved large language models (LLMs) by allowing them to decompose problems into intermediate steps. While CoT is widely effectiv…
DigitalCoach: Communication and Grounding Gaps in Human and Agentic Computer Use Coaching
Meng Chen, Anya Ji, Tsung-Han Wu +4
Agents are increasingly capable of automating software tasks, but can they teach humans how to use software themselves? We introduce DigitalCoach, a multimodal dataset of 72 human…
TopBench: A Benchmark for Implicit Predictive Reasoning in Tabular Question Answering
An-Yang Ji, Jun-Peng Jiang, De-Chuan Zhan +1
Large Language Models (LLMs) have advanced Table Question Answering, where most queries can be answered by extracting information or simple aggregation. However, a common class of…
Ad hoc conventions generalize to new referents
Anya Ji, Claire Augusta Bergey, Ron Eliav +2
How do people talk about things they've never talked about before? One view suggests that a new shared naming system establishes an arbitrary link to a specific target, like proper…
Semantic uncertainty guides the extension of conventions to new referents
Ron Eliav, Anya Ji, Yoav Artzi +1
A long tradition of studies in psycholinguistics has examined the formation and generalization of ad hoc conventions in reference games, showing how newly acquired conventions for…
Abstract Visual Reasoning with Tangram Shapes
Anya Ji, Noriyuki Kojima, Noah Rush +4
We introduce KiloGram, a resource for studying abstract visual reasoning in humans and machines. Drawing on the history of tangram puzzles as stimuli in cognitive science, we build…