activity
20242026
collaborators

5 papers

cs.CV2026

Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answering

Yura Choi, Roy Miles, Rolandos Alexandros Potamias +3

Understanding and answering questions based on a user's pointing gesture is essential for next-generation egocentric AI assistants. However, current Multimodal Large Language Model…

cs.CV2025

SATGround: A Spatially-Aware Approach for Visual Grounding in Remote Sensing

Aysim Toker, Andreea-Maria Oncescu, Roy Miles +2

Vision-language models (VLMs) are emerging as powerful generalist tools for remote sensing, capable of integrating information across diverse tasks and enabling flexible, instructi…

cs.CV2025

RetouchLLM: Training-free Code-based Image Retouching with Vision Language Models

Moon Ye-Bin, Roy Miles, Tae-Hyun Oh +2

Image retouching not only enhances visual quality but also serves as a means of expressing personal preferences and emotions. However, existing learning-based approaches require la…

cs.CV2025

"Principal Components" Enable A New Language of Images

Xin Wen, Bingchen Zhao, Ismail Elezi +2

We introduce a novel visual tokenization framework that embeds a provable PCA-like structure into the latent token space. While existing visual tokenizers primarily optimize for re…

cs.CL2024

From Attention to Activation: Unravelling the Enigmas of Large Language Models

Prannay Kaul, Chengcheng Ma, Ismail Elezi +1

We study two strange phenomena in auto-regressive Transformers: (1) the dominance of the first token in attention heads; (2) the occurrence of large outlier activations in the hidd…