1 citations · 1 across the 14 of their papers we have counts for
12 papers · 1 filter
EquiSteer: Cross-Attention Steering Towards a Fairer Text-Guided Image Generation
Tatiana Gaintseva, Akshit Achara, Gregory Slabaugh +2
Text-to-image diffusion models power everyday creative tasks, but they still reproduce the demographic biases in their training data. On common prompts such as ``a photo of a nurse…
Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answering
Yura Choi, Roy Miles, Rolandos Alexandros Potamias +3
Understanding and answering questions based on a user's pointing gesture is essential for next-generation egocentric AI assistants. However, current Multimodal Large Language Model…
DreamCAD: Scaling Multi-modal CAD Generation using Differentiable Parametric Surfaces
Mohammad Sadil Khan, Muhammad Usama, Rolandos Alexandros Potamias +4
Computer-Aided Design (CAD) relies on structured and editable geometric representations, yet existing generative methods are constrained by small annotated datasets with explicit d…
SATGround: A Spatially-Aware Approach for Visual Grounding in Remote Sensing
Aysim Toker, Andreea-Maria Oncescu, Roy Miles +2
Vision-language models (VLMs) are emerging as powerful generalist tools for remote sensing, capable of integrating information across diverse tasks and enabling flexible, instructi…
RetouchLLM: Training-free Code-based Image Retouching with Vision Language Models
Moon Ye-Bin, Roy Miles, Tae-Hyun Oh +2
Image retouching not only enhances visual quality but also serves as a means of expressing personal preferences and emotions. However, existing learning-based approaches require la…
Region-based Cluster Discrimination for Visual Representation Learning
Yin Xie, Kaicheng Yang, Xiang An +9
Learning visual representations is foundational for a broad spectrum of downstream tasks. Although recent vision-language contrastive models, such as CLIP and SigLIP, have achieved…