7 papers
cotomi Act: Learning to Automate Work by Watching You
Masafumi Oyamada, Kunihiro Takeoka, Kosuke Akimoto +5
What if a browser agent could learn your work simply by watching you do it? We present cotomi Act, a browser-based computer-using agent that combines reliable multi-step task execu…
Shape vs. Context: Examining Human--AI Gaps in Ambiguous Japanese Character Recognition
Daichi Haraguchi
High text recognition performance does not guarantee that Vision-Language Models (VLMs) share human-like decision patterns when resolving ambiguity. We investigate this behavioral…
Automatic Text Box Placement for Supporting Typographic Design
Jun Muraoka, Daichi Haraguchi, Naoto Inoue +3
In layout design for advertisements and web pages, balancing visual appeal and communication efficiency is crucial. This study examines automated text box placement in incomplete l…
Few-Part-Shot Font Generation
Masaki Akiba, Shumpei Takezaki, Daichi Haraguchi +1
This paper proposes a novel model of few-part-shot font generation, which designs an entire font based on a set of partial design elements, i.e., partial shapes. Unlike conventiona…
Total Disentanglement of Font Images into Style and Character Class Features
Daichi Haraguchi, Wataru Shimoda, Kota Yamaguchi +1
In this paper, we demonstrate a total disentanglement of font images. Total disentanglement is a neural network-based method for decomposing each font image nonlinearly and complet…
MG-Gen: Single Image to Motion Graphics Generation
Takahiro Shirakawa, Tomoyuki Suzuki, Takuto Narumoto +1
We introduce MG-Gen, a framework that generates motion graphics directly from a single raster image. MG-Gen decompose a single raster image into layered structures represented as H…