10 citations · 25 across the 10 of their papers we have counts for
14 papers
Semantic Document Derendering: SVG Reconstruction via Vision-Language Modeling
Adam Hazimeh, Ke Wang, Mark Collier +3
Multimedia documents such as slide presentations and posters are designed to be interactive and easy to modify. Yet, they are often distributed in a static raster format, which lim…
Sketch-to-Layout: Sketch-Guided Multimodal Layout Generation
Riccardo Brioschi, Aleksandr Alekseev, Emanuele Nevali +9
Graphic layout generation is a growing research area focusing on generating aesthetically pleasing layouts ranging from poster designs to documents. While recent research has explo…
Representing Online Handwriting for Recognition in Large Vision-Language Models
Anastasiia Fadeeva, Philippe Schlattner, Andrii Maksai +4
The adoption of tablets with touchscreens and styluses is increasing, and a key feature is converting handwriting to text, enabling search, indexing, and AI assistance. Meanwhile,…
Pi-DUAL: Using Privileged Information to Distinguish Clean from Noisy Labels
Ke Wang, Guillermo Ortiz-Jimenez, Rodolphe Jenatton +3
Label noise is a pervasive problem in deep learning that often compromises the generalization performance of trained models. Recently, leveraging privileged information (PI) -- inf…
Three Towers: Flexible Contrastive Learning with Pretrained Image Models
Jannik Kossen, Mark Collier, Basil Mustafa +7
We introduce Three Towers (3T), a flexible method to improve the contrastive learning of vision-language models by incorporating pretrained image classifiers. While contrastive mod…
SmartChoices: Augmenting Software with Learned Implementations
Daniel Golovin, Gabor Bartok, Eric Chen +7
In many software systems, heuristics are used to make decisions - such as cache eviction, task scheduling, and information presentation - that have a significant impact on overall…