9 papers
Walk and Read Less: Improving the Efficiency of Vision-and-Language Navigation via Tuning-Free Multimodal Token Pruning
Wenda Qin, Andrea Burns, Bryan A. Plummer +1
Large models achieve strong performance on Vision-and-Language Navigation (VLN) tasks, but are costly to run in resource-limited environments. Token pruning offers appealing tradeo…
Gen-AFFECT: Generation of Avatar Fine-grained Facial Expressions with Consistent identiTy
Hao Yu, Rupayan Mallick, Margrit Betke +1
Different forms of customized 2D avatars are widely used in gaming applications, virtual communication, education, and content creation. However, existing approaches often fail to…
LoRA-Loop: Closing the Synthetic Replay Cycle for Continual VLM Learning
Kaihong Wang, Donghyun Kim, Margrit Betke
Continual learning for vision-language models has achieved remarkable performance through synthetic replay, where samples are generated using Stable Diffusion to regularize during…
ExeChecker: Where Did I Go Wrong?
Yiwen Gu, Mahir Patel, Margrit Betke
In this paper, we present a contrastive learning based framework, ExeChecker, for the interpretation of rehabilitation exercises. Our work builds upon state-of-the-art advances in…
GenEAva: Generating Cartoon Avatars with Fine-Grained Facial Expressions from Realistic Diffusion-based Faces
Hao Yu, Rupayan Mallick, Margrit Betke +1
Cartoon avatars have been widely used in various applications, including social media, online tutoring, and gaming. However, existing cartoon avatar datasets and generation methods…
DebiasPI: Inference-time Debiasing by Prompt Iteration of a Text-to-Image Generative Model
Sarah Bonna, Yu-Cheng Huang, Ekaterina Novozhilova +11
Ethical intervention prompting has emerged as a tool to counter demographic biases of text-to-image generative AI models. Existing solutions either require to retrain the model or…