6 papers · 1 filter
Think, Look, and Revise: Inconsistency-Aware Visual Self-Correction in MLLMs
Yu Cheng, Arushi Goel, Hakan Bilen
Tool-augmented multimodal reasoning integrates external tools (e.g., object detection, depth estimation) into multimodal large language models (MLLMs) to address perceptual bottlen…
Beyond Pixel Histories: World Models with Persistent 3D State
Samuel Garcin, Thomas Walker, Steven McDonagh +5
Interactive world models continually generate video by responding to a user's actions, enabling open-ended generation capabilities. However, existing models typically lack a 3D rep…
Enabling Progressive Whole-slide Image Analysis with Multi-scale Pyramidal Network
Shuyang Wu, Yifu Qiu, Ines P Nearchou +4
Multiple-instance Learning (MIL) is commonly used for computational pathology (CPath), where multi-scale features are essential for capturing both fine cellular details and broad t…
Visually Interpretable Subtask Reasoning for Visual Question Answering
Yu Cheng, Arushi Goel, Hakan Bilen
Answering complex visual questions like `Which red furniture can be used for sitting?' requires multi-step reasoning, including object recognition, attribute filtering, and relatio…
Coarse or Fine? Recognising Action End States without Labels
Davide Moltisanti, Hakan Bilen, Laura Sevilla-Lara +1
We focus on the problem of recognising the end state of an action in an image, which is critical for understanding what action is performed and in which manner. We study this focus…
Multi-task Learning with 3D-Aware Regularization
Wei-Hong Li, Steven McDonagh, Ales Leonardis +1
Deep neural networks have become a standard building block for designing models that can perform multiple dense computer vision tasks such as depth estimation and semantic segmenta…