collaborators

8 papers

cs.CV2025

FraQAT: Quantization Aware Training with Fractional bits

Luca Morreale, Alberto Gil C. P. Ramos, Malcolm Chadwick +4

State-of-the-art (SOTA) generative models have demonstrated impressive capabilities in image synthesis or text generation, often with a large capacity model. However, these large m…

cs.CV2025

Efficient High-Resolution Image Editing with Hallucination-Aware Loss and Adaptive Tiling

Young D. Kwon, Abhinav Mehrotra, Malcolm Chadwick +2

High-resolution (4K) image-to-image synthesis has become increasingly important for mobile applications. Existing diffusion models for image editing face significant challenges, in…

eess.AS2025

Robust Unsupervised Adaptation of a Speech Recogniser Using Entropy Minimisation and Speaker Codes

Rogier C. van Dalen, Shucong Zhang, Titouan Parcollet +1

Speech recognisers usually perform optimally only in a specific environment and need to be adapted to work well in another. For adaptation to a new speaker, there is often too litt…

eess.AS2025

Evaluation of LLMs in Speech is Often Flawed: Test Set Contamination in Large Language Models for Speech Recognition

Yuan Tseng, Titouan Parcollet, Rogier van Dalen +2

Recent work suggests that large language models (LLMs) can improve performance of speech tasks compared to existing systems. To support their claims, results on LibriSpeech and Com…

cs.CV2025

Guidance Free Image Editing via Explicit Conditioning

Mehdi Noroozi, Alberto Gil Ramos, Luca Morreale +4

Current sampling mechanisms for conditional diffusion models rely mainly on Classifier Free Guidance (CFG) to generate high-quality images. However, CFG requires several denoising…

cs.CV2025

Upcycling Text-to-Image Diffusion Models for Multi-Task Capabilities

Ruchika Chavhan, Abhinav Mehrotra, Malcolm Chadwick +4

Text-to-image synthesis has witnessed remarkable advancements in recent years. Many attempts have been made to adopt text-to-image models to support multiple tasks. However, existi…