activity
20182026
most citedLike Article, Like Audience: Enforcing Multimodal Correlations for Disinformation Detection

31 citations · 77 across the 49 of their papers we have counts for

collaborators
Showing cs.CVShow all

19 papers · 1 filter

cs.CV2026

What Makes Linguistic Representations Good Models of High-Level Visual Perception in the Human Brain?

Anna Bavaresco, Ina Klarić, Raquel Fernández +1

Image descriptions represented with language models (LMs) predict human brain responses to naturalistic images in high-level visual regions, but the factors driving this predictivi…

cs.CV2025

Automated Prompt Generation for Creative and Counterfactual Text-to-image Synthesis

Aleksa Jelaca, Ying Jiao, Chang Tian +1

Text-to-image generation has advanced rapidly with large-scale multimodal training, yet fine-grained controllability remains a critical challenge. Counterfactual controllability, d…

cs.CV2025

Consistent Story Generation: Unlocking the Potential of Zigzag Sampling

Mingxiao Li, Mang Ning, Marie-Francine Moens

Text-to-image generation models have made significant progress in producing high-quality images from textual descriptions, yet they continue to struggle with maintaining subject co…

cs.CV2024

Action-based image editing guided by human instructions

Maria Mihaela Trusca, Mingxiao Li, Marie-Francine Moens

Text-based image editing is typically approached as a static task that involves operations such as inserting, deleting, or modifying elements of an input image based on human instr…

cs.CV2024

DistilDoc: Knowledge Distillation for Visually-Rich Document Applications

Jordy Van Landeghem, Subhajit Maity, Ayan Banerjee +4

This work explores knowledge distillation (KD) for visually-rich document (VRD) applications such as document layout analysis (DLA) and document image classification (DIC). While V…

cs.CV2024

Multimodal Adaptive Inference for Document Image Classification with Anytime Early Exiting

Omar Hamed, Souhail Bakkali, Marie-Francine Moens +2

This work addresses the need for a balanced approach between performance and efficiency in scalable production environments for visually-rich document understanding (VDU) tasks. Cu…