activity
20242026
most citedGLOV: Guided Large Language Models as Implicit Optimizers for Vision Language Models

1 citations · 1 across the 10 of their papers we have counts for

collaborators
Showing 2025Show all

5 papers · 1 filter

cs.CV2025

PRISMM-Bench: A Benchmark of Peer-Review Grounded Multimodal Inconsistencies

Lukas Selch, Yufang Hou, M. Jehanzeb Mirza +4

Large Multimodal Models (LMMs) are increasingly applied to scientific research, yet it remains unclear whether they can reliably understand and reason over the multimodal complexit…

cs.CV2025

TTRV: Test-Time Reinforcement Learning for Vision Language Models

Akshit Singh, Shyam Marjit, Wei Lin +7

Existing methods for extracting reward signals in Reinforcement Learning typically rely on labeled data and dedicated training splits, a setup that contrasts with how humans learn…

cs.CV2025

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes

Paul Gavrikov, Wei Lin, M. Jehanzeb Mirza +6

Is basic visual understanding really solved in state-of-the-art VLMs? We present VisualOverload, a slightly different visual question answering (VQA) benchmark comprising 2,720 que…

cs.CV2025

Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion

Jacob Hansen, Wei Lin, Junmo Kang +6

Visual Instruction Tuning (VisIT) data, commonly available as human-assistant conversations with images interleaved in the human turns, are currently the most widespread vehicle fo…

cs.CV2025

Exploring Modality Guidance to Enhance VFM-based Feature Fusion for UDA in 3D Semantic Segmentation

Johannes Spoecklberger, Wei Lin, Pedro Hermosilla +3

Vision Foundation Models (VFMs) have become a de facto choice for many downstream vision tasks, like image classification, image segmentation, and object localization. However, the…