4 papers
A Scalable Multi-Task Model for Virtual Sensors
Leon Götz, Lars Frederik Peiss, Erik Sauer +4
Virtual sensors replace expensive physical sensors in critical applications through machine learning by predicting target signals from available measurements. Existing virtual sens…
Unexplored flaws in multiple-choice VQA evaluations
Fabio Rosenthal, Sebastian Schmidt, Thorsten Graf +3
Multimodal Large Language Models (MLLMs) demonstrate strong capabilities in handling image-text inputs. A common way to assess this ability is through multiple-choice Visual Questi…
FOCUS: Internal MLLM Representations for Efficient Fine-Grained Visual Question Answering
Liangyu Zhong, Fabio Rosenthal, Joachim Sicking +4
While Multimodal Large Language Models (MLLMs) offer strong perception and reasoning capabilities for image-text input, Visual Question Answering (VQA) focusing on small image deta…
Investigating the Robustness of Retrieval-Augmented Generation at the Query Level
Sezen Perçin, Xin Su, Qutub Sha Syed +4
Large language models (LLMs) are very costly and inefficient to update with new information. To address this limitation, retrieval-augmented generation (RAG) has been proposed as a…