activity
20232026
collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

Think, Look, and Revise: Inconsistency-Aware Visual Self-Correction in MLLMs

Yu Cheng, Arushi Goel, Hakan Bilen

Tool-augmented multimodal reasoning integrates external tools (e.g., object detection, depth estimation) into multimodal large language models (MLLMs) to address perceptual bottlen…

cs.CV2026

Beyond Pixel Histories: World Models with Persistent 3D State

Samuel Garcin, Thomas Walker, Steven McDonagh +5

Interactive world models continually generate video by responding to a user's actions, enabling open-ended generation capabilities. However, existing models typically lack a 3D rep…

cs.CV2026

Enabling Progressive Whole-slide Image Analysis with Multi-scale Pyramidal Network

Shuyang Wu, Yifu Qiu, Ines P Nearchou +4

Multiple-instance Learning (MIL) is commonly used for computational pathology (CPath), where multi-scale features are essential for capturing both fine cellular details and broad t…

cs.CV2025

Visually Interpretable Subtask Reasoning for Visual Question Answering

Yu Cheng, Arushi Goel, Hakan Bilen

Answering complex visual questions like `Which red furniture can be used for sitting?' requires multi-step reasoning, including object recognition, attribute filtering, and relatio…

cs.CV2024

Coarse or Fine? Recognising Action End States without Labels

Davide Moltisanti, Hakan Bilen, Laura Sevilla-Lara +1

We focus on the problem of recognising the end state of an action in an image, which is critical for understanding what action is performed and in which manner. We study this focus…

cs.CV2023

Multi-task Learning with 3D-Aware Regularization

Wei-Hong Li, Steven McDonagh, Ales Leonardis +1

Deep neural networks have become a standard building block for designing models that can perform multiple dense computer vision tasks such as depth estimation and semantic segmenta…