31 citations · 77 across the 49 of their papers we have counts for
19 papers · 1 filter
What Makes Linguistic Representations Good Models of High-Level Visual Perception in the Human Brain?
Anna Bavaresco, Ina Klarić, Raquel Fernández +1
Image descriptions represented with language models (LMs) predict human brain responses to naturalistic images in high-level visual regions, but the factors driving this predictivi…
Automated Prompt Generation for Creative and Counterfactual Text-to-image Synthesis
Aleksa Jelaca, Ying Jiao, Chang Tian +1
Text-to-image generation has advanced rapidly with large-scale multimodal training, yet fine-grained controllability remains a critical challenge. Counterfactual controllability, d…
Consistent Story Generation: Unlocking the Potential of Zigzag Sampling
Mingxiao Li, Mang Ning, Marie-Francine Moens
Text-to-image generation models have made significant progress in producing high-quality images from textual descriptions, yet they continue to struggle with maintaining subject co…
Action-based image editing guided by human instructions
Maria Mihaela Trusca, Mingxiao Li, Marie-Francine Moens
Text-based image editing is typically approached as a static task that involves operations such as inserting, deleting, or modifying elements of an input image based on human instr…
DistilDoc: Knowledge Distillation for Visually-Rich Document Applications
Jordy Van Landeghem, Subhajit Maity, Ayan Banerjee +4
This work explores knowledge distillation (KD) for visually-rich document (VRD) applications such as document layout analysis (DLA) and document image classification (DIC). While V…
Multimodal Adaptive Inference for Document Image Classification with Anytime Early Exiting
Omar Hamed, Souhail Bakkali, Marie-Francine Moens +2
This work addresses the need for a balanced approach between performance and efficiency in scalable production environments for visually-rich document understanding (VDU) tasks. Cu…