From the 1 of 6 linked papers with an AI index.
6 papers
MED-DSLC: Multi-Expert-Domain Classification via Domain Supervision and Logit Calibration
Zheng Zeng, Deepak Sridhar, Nuno Vasconcelos
The paper introduces MED-DSLC, a lightweight method that restores global logit comparability across independently fine‑tuned vision‑language model experts by using domain supervisi…
Video Reasoning without Training
Deepak Sridhar, Kartikeya Bhardwaj, Jeya Pradha Jeyaraj +3
Video reasoning using Large Multimodal Models (LMMs) relies on costly reinforcement learning (RL) and verbose chain-of-thought, resulting in substantial computational overhead duri…
Diffusion Models with Adaptive Negative Sampling Without External Resources
Alakh Desai, Nuno Vasconcelos
Diffusion models (DMs) have demonstrated an unparalleled ability to create diverse and high-fidelity images from text prompts. However, they are also well-known to vary substantial…
SCHEME: Scalable Channel Mixer for Vision Transformers
Deepak Sridhar, Yunsheng Li, Nuno Vasconcelos
Vision Transformers have achieved impressive performance in many vision tasks. While the token mixer or attention block has been studied in great detail, much less research has bee…
Improving image synthesis with diffusion-negative sampling
Alakh Desai, Nuno Vasconcelos
For image generation with diffusion models (DMs), a negative prompt n can be used to complement the text prompt p, helping define properties not desired in the synthesized image. W…
Adapting Diffusion Models for Improved Prompt Compliance and Controllable Image Synthesis
Deepak Sridhar, Abhishek Peri, Rohith Rachala +1
Recent advances in generative modeling with diffusion processes (DPs) enabled breakthroughs in image synthesis. Despite impressive image quality, these models have various prompt c…