papers

Publications (33)

cs.CV2024

Concept Weaver: Enabling Multi-Concept Fusion in Text-to-Image Models

Gihyun Kwon, Simon Jenni, Dingzeyu Li +3

While there has been significant progress in customizing text-to-image generation models, generating images that combine multiple personalized concepts remains challenging. In this…

cs.CV2023

Spatio-Temporal Crop Aggregation for Video Representation Learning

Sepehr Sameni, Simon Jenni, Paolo Favaro

We propose Spatio-temporal Crop Aggregation for video representation LEarning (SCALE), a novel method that enjoys high scalability at both training and inference time. Our model bu…

cs.CV2020

Video Representation Learning by Recognizing Temporal Transformations

Simon Jenni, Givi Meishvili, Paolo Favaro

We introduce a novel self-supervised learning approach to learn representations of videos that are responsive to changes in the motion dynamics. Our representations can be learned…

cs.CV2023

Audio-Visual Contrastive Learning with Temporal Self-Supervision

Simon Jenni, Alexander Black, John Collomosse

We propose a self-supervised learning approach for videos that learns representations of both the RGB frames and the accompanying audio without human supervision. In contrast to im…

cs.CV2023

Meta-Personalizing Vision-Language Models to Find Named Instances in Video

Chun-Hsiao Yeh, Bryan Russell, Josef Sivic +2

Large-scale vision-language models (VLM) have shown impressive results for language-guided search applications. While these models allow category-level queries, they currently stru…

cs.CV2024

Sync from the Sea: Retrieving Alignable Videos from Large-Scale Datasets

Ishan Rajendrakumar Dave, Fabian Caba Heilbron, Mubarak Shah +1

Temporal video alignment aims to synchronize the key events like object interactions or action phase transitions in two videos. Such methods could benefit various video editing, pr…

cs.CV2021

VPN: Video Provenance Network for Robust Content Attribution

Alexander Black, Tu Bui, Simon Jenni +2

We present VPN - a content attribution method for recovering provenance information from videos shared online. Platforms, and users, often transform video into different quality, c…

cs.CV2026

Seeing Through Words: Controlling Visual Retrieval Quality with Language Models

Jianglin Lu, Simon Jenni, Kushal Kafle +3

Text-to-image retrieval is a fundamental task in vision-language learning, yet in real-world scenarios it is often challenged by short and underspecified user queries. Such queries…

cs.CV2020

Learning to Have an Ear for Face Super-Resolution

Givi Meishvili, Simon Jenni, Paolo Favaro

We propose a novel method to use both audio and a low-resolution image to perform extreme face super-resolution (a 16x increase of the input size). When the resolution of the input…

cs.CL2025

MAGNET: Augmenting Generative Decoders with Representation Learning and Infilling Capabilities

Savya Khosla, Aditi Tiwari, Kushal Kafle +4

While originally designed for unidirectional generative modeling, decoder-only large language models (LLMs) are increasingly being adapted for bidirectional modeling. However, unid…

cs.CV2025

More Than the Final Answer: Improving Visual Extraction and Logical Consistency in Vision-Language Models

Hoang Anh Just, Yifei Fan, Handong Zhao +6

Reinforcement learning from verifiable rewards (RLVR) has recently been extended from text-only LLMs to vision-language models (VLMs) to elicit long-chain multimodal reasoning. How…

cs.CV2025

The Photographer Eye: Teaching Multimodal Large Language Models to Understand Image Aesthetics like Photographers

Daiqing Qi, Handong Zhao, Jing Shi +5

While editing directly from life, photographers have found it too difficult to see simultaneously both the blue and the sky. Photographer and curator, Szarkowski insightfully revea…

cs.CV2020

Steering Self-Supervised Feature Learning Beyond Local Pixel Statistics

Simon Jenni, Hailin Jin, Paolo Favaro

We introduce a novel principle for self-supervised feature learning based on the discrimination of specific transformations of an image. We argue that the generalization capability…

cs.CV2026

RetouchIQ: MLLM Agents for Instruction-Based Image Retouching with Generalist Reward

Qiucheng Wu, Jing Shi, Simon Jenni +4

Recent advances in multimodal large language models (MLLMs) have shown great potential for extending vision-language reasoning to professional tool-based image editing, enabling in…

cs.CV2025

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory

Sethuraman TV, Savya Khosla, Vignesh Srinivasakumar +5

Dense video prediction tasks, such as object tracking and semantic segmentation, require video encoders that generate temporally consistent, spatially dense features for every fram…

cs.CR2023

DECORAIT -- DECentralized Opt-in/out Registry for AI Training

Kar Balan, Alex Black, Simon Jenni +3

We present DECORAIT; a decentralized registry through which content creators may assert their right to opt in or out of AI training as well as receive reward for their contribution…

cs.CV2023

Representation Learning by Detecting Incorrect Location Embeddings

Sepehr Sameni, Simon Jenni, Paolo Favaro

In this paper, we introduce a novel self-supervised learning (SSL) loss for image representation learning. There is a growing belief that generalization in deep neural networks is…

cs.CV2023

VADER: Video Alignment Differencing and Retrieval

Alexander Black, Simon Jenni, Tu Bui +5

We propose VADER, a spatio-temporal matching, alignment, and change summarization method to help fight misinformation spread via manipulated videos. VADER matches and coarsely alig…

cs.CV2026

Gen2Balance: Generative Balancing for Long-Tailed Video Action Recognition

Prajwal Gatti, Simon Jenni, Fabian Caba Heilbron +1

We address the problem of training on long-tailed data for video action recognition. We propose to augment the training set using a text-to-video generative model, conditioned on d…

cs.CV2019

On Stabilizing Generative Adversarial Training with Noise

Simon Jenni, Paolo Favaro

We present a novel method and analysis to train generative adversarial networks (GAN) in a stable manner. As shown in recent analysis, training is often undermined by the probabili…

cs.CV2026

Stress Tests REVEAL Fragile Temporal and Visual Grounding in Video-Language Models

Sethuraman T, Savya Khosla, Aditi Tiwari +11

This work investigates a fundamental question: Do Video-Language Models (VidLMs) robustly account for video content, temporal sequence, and motion? Our investigation shows that, su…

cs.CV2021

Time-Equivariant Contrastive Video Representation Learning

Simon Jenni, Hailin Jin

We introduce a novel self-supervised contrastive learning method to learn representations from unlabelled videos. Existing approaches ignore the specifics of input distortions, e.g…

cs.CV2023

SImProv: Scalable Image Provenance Framework for Robust Content Attribution

Alexander Black, Tu Bui, Simon Jenni +3

We present SImProv - a scalable image provenance framework to match a query image back to a trusted database of originals and identify possible manipulations on the query. SImProv…

cs.CV2020

Self-Supervised Multi-View Synchronization Learning for 3D Pose Estimation

Simon Jenni, Paolo Favaro

Current state-of-the-art methods cast monocular 3D human pose estimation as a learning problem by training neural networks on large data sets of images and corresponding skeleton p…

cs.CV2023

No More Shortcuts: Realizing the Potential of Temporal Self-Supervision

Ishan Rajendrakumar Dave, Simon Jenni, Mubarak Shah

Self-supervised approaches for video have shown impressive results in video understanding tasks. However, unlike early works that leverage temporal self-supervision, current state-…

cs.CV2026

The Indra Representation Hypothesis for Multimodal Alignment

Jianglin Lu, Hailing Wang, Kuo Yang +3

Recent studies have uncovered an interesting phenomenon: unimodal foundation models tend to learn convergent representations, regardless of differences in architecture, training ob…

cs.CV2021

Learning to Deblur and Rotate Motion-Blurred Faces

Givi Meishvili, Attila Szabó, Simon Jenni +1

We propose a solution to the novel task of rendering sharp videos from new viewpoints from a single motion-blurred image of a face. Our method handles the complexity of face blur b…

cs.CV2024

FINEMATCH: Aspect-based Fine-grained Image and Text Mismatch Detection and Correction

Hang Hua, Jing Shi, Kushal Kafle +5

Recent progress in large-scale pre-training has led to the development of advanced vision-language models (VLMs) with remarkable proficiency in comprehending and generating multimo…

cs.CV2018

Self-Supervised Feature Learning by Learning to Spot Artifacts

Simon Jenni, Paolo Favaro

We introduce a novel self-supervised learning method based on adversarial training. Our objective is to train a discriminator network to distinguish real images from images with sy…

cs.CV2018

Deep Bilevel Learning

Simon Jenni, Paolo Favaro

We present a novel regularization approach to train neural networks that enjoys better generalization and test error than standard stochastic gradient descent. Our approach is base…

cs.CV2022

Video-ReTime: Learning Temporally Varying Speediness for Time Remapping

Simon Jenni, Markus Woodson, Fabian Caba Heilbron

We propose a method for generating a temporally remapped video that matches the desired target duration while maximally preserving natural video dynamics. Our approach trains a neu…

cs.CV2025

Improving Large Vision and Language Models by Learning from a Panel of Peers

Jefferson Hernandez, Jing Shi, Simon Jenni +2

Traditional alignment methods for Large Vision and Language Models (LVLMs) primarily rely on human-curated preference data. Human-generated preference data is costly; machine-gener…

cs.CV2023

EKILA: Synthetic Media Provenance and Attribution for Generative Art

Kar Balan, Shruti Agarwal, Simon Jenni +3

We present EKILA; a decentralized framework that enables creatives to receive recognition and reward for their contributions to generative AI (GenAI). EKILA proposes a robust visua…