activity
20162024
most citedDoes CLIP Benefit Visual Question Answering in the Medical Domain as Much as it Does in the General Domain?

41 citations · 58 across the 7 of their papers we have counts for

collaborators

9 papers

cs.CL2024

Language Adaptation on a Tight Academic Compute Budget: Tokenizer Swapping Works and Pure bfloat16 Is Enough

Konstantin Dobler, Gerard de Melo

We investigate continued pretraining of LLMs for language adaptation on a tight academic budget: a setting in which only a few GPUs can be used in parallel, for a heavily constrain…

cs.CR2024

MaskAnyone Toolkit: Offering Strategies for Minimizing Privacy Risks and Maximizing Utility in Audio-Visual Data Archiving

Babajide Alamu Owoyele, Martin Schilling, Rohan Sawahn +5

This paper introduces MaskAnyone, a novel toolkit designed to navigate some privacy and ethical concerns of sharing audio-visual data in research. MaskAnyone offers a scalable, use…

cs.CL2024

CommitBench: A Benchmark for Commit Message Generation

Maximilian Schall, Tamara Czinczoll, Gerard de Melo

Writing commit messages is a tedious daily task for many software developers, and often remains neglected. Automating this task has the potential to save time while ensuring that m…

cs.AI2023

IntentDial: An Intent Graph based Multi-Turn Dialogue System with Reasoning Path Visualization

Zengguang Hao, Jie Zhang, Binxia Xu +3

Intent detection and identification from multi-turn dialogue has become a widely explored technique in conversational agents, for example, voice assistants and intelligent customer…

cs.CV2023

FARSEC: A Reproducible Framework for Automatic Real-Time Vehicle Speed Estimation Using Traffic Cameras

Lucas Liebe, Franz Sauerwald, Sylwester Sawicki +6

Estimating the speed of vehicles using traffic cameras is a crucial task for traffic surveillance and management, enabling more optimal traffic flow, improved road safety, and lowe…

cs.CV2023

MultiModal Bias: Introducing a Framework for Stereotypical Bias Assessment beyond Gender and Race in Vision Language Models

Sepehr Janghorbani, Gerard de Melo

Recent breakthroughs in self supervised training have led to a new class of pretrained vision language models. While there have been investigations of bias in multimodal models, th…