From the 1 of 7 linked papers with an AI index.
7 papers
OvisOCR2 Technical Report
Shiyin Lu, Yinglun Li, Yu Xia +10
OvisOCR2 is a 0.8 B parameter end‑to‑end model that converts document page images into Markdown, handling text, formulas, tables, and visual regions, and achieves state‑of‑the‑art…
DigitalCoach: Communication and Grounding Gaps in Human and Agentic Computer Use Coaching
Meng Chen, Anya Ji, Tsung-Han Wu +4
Agents are increasingly capable of automating software tasks, but can they teach humans how to use software themselves? We introduce DigitalCoach, a multimodal dataset of 72 human…
Animation2Code: Evaluating Temporal Visual Reasoning in Video-to-Code Generation
Anya Ji, Abhijith Varma Mudunuri, David M. Chan +1
While recent vision-language models (VLMs) have achieved significant improvements on static visual-to-code tasks such as generating code for webpages, charts, or SVGs, it remains u…
TopBench: A Benchmark for Implicit Predictive Reasoning in Tabular Question Answering
An-Yang Ji, Jun-Peng Jiang, De-Chuan Zhan +1
Large Language Models (LLMs) have advanced Table Question Answering, where most queries can be answered by extracting information or simple aggregation. However, a common class of…
ScribbleEdit: Synthetic Data for Image Editing with Scribbles and Text
Anya Ji, George Ma, Téa Wright +4
Recent progress in generative models has significantly advanced image editing capabilities, yet precise and intuitive user control remains difficult. Specifically, users often stru…
Ad hoc conventions generalize to new referents
Anya Ji, Claire Augusta Bergey, Ron Eliav +2
How do people talk about things they've never talked about before? One view suggests that a new shared naming system establishes an arbitrary link to a specific target, like proper…